vix.ing · top · new · best · stats · spec

Privacy- and Utility-Preserving Textual Analysis via Calibrated\n Multivariate Perturbations

2019/10/20 by Oluwaseyi Feyisetan, Feyisetan, Oluwaseyi, Borja Balle +5 · 10 citations
Computer Science · #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data

paper · pdf · doi:10.48550/arxiv.1910.08902

openalex publication_date 2019/10/20 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28

Abstract

Accurately learning from user data while providing quantifiable privacy\nguarantees provides an opportunity to build better ML models while maintaining\nuser trust. This paper presents a formal approach to carrying out privacy\npreserving text perturbation using the notion of dx-privacy designed to achieve\ngeo-indistinguishability in location data. Our approach applies carefully\ncalibrated noise to vector representation of words in a high dimension space as\ndefined by word embedding models. We present a privacy proof that satisfies\ndx-privacy where the privacy parameter epsilon provides guarantees with respect\nto a distance metric defined by the word embedding space. We demonstrate how\nepsilon can be selected by analyzing plausible deniability statistics backed up\nby large scale analysis on GloVe and fastText embeddings. We conduct privacy\naudit experiments against 2 baseline models and utility experiments on 3\ndatasets to demonstrate the tradeoff between privacy and utility for varying\nvalues of epsilon on different task types. Our results demonstrate practical\nutility (< 2% utility loss for training binary classifiers) while providing\nbetter privacy guarantees than baseline models.\n

Cited by

Related