vix.ing · top · new · best · stats

Adversarial representation learning for private speech generation

2020/06/16 by David Ericsson, Adam Östberg, Ericsson, David +7 · 14 citations
Computer Science · Engineering · #Adversarial system #Artificial intelligence #Audio and Speech Processing (eess.AS) #Benchmark (surveying) #Computer science #Computer vision #Convolutional neural network #Domain (mathematical analysis) #FOS: Computer and information sciences #FOS: Electrical engineering #Filter (signal processing) #Generative Adversarial Networks and Image Synthesis #Generative grammar #Machine Learning (cs.LG) #Music and Audio Processing #Natural language processing #Raw data #Representation (politics) #Sound (cs.SD) #Spectrogram #Speech Recognition and Synthesis #Speech recognition #Utterance #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2006.09114

published in arXiv (Cornell University) (Cornell University) · Submitted to ICML 2020 Workshop on Self-supervision in Audio and Speech (SAS)

openalex publication_date 2020/06/16 · arxiv created 2020/06/17 · arxiv updated 2020/06/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

As more and more data is collected in various settings across organizations, companies, and countries, there has been an increase in the demand of user privacy. Developing privacy preserving methods for data analytics is thus an important area of research. In this work we present a model based on generative adversarial networks (GANs) that learns to obfuscate specific sensitive attributes in speech data. We train a model that learns to hide sensitive information in the data, while preserving the meaning in the utterance. The model is trained in two steps: first to filter sensitive information in the spectrogram domain, and then to generate new and private information independent of the filtered one. The model is based on a U-Net CNN that takes mel-spectrograms as input. A MelGAN is used to invert the spectrograms back to raw audio waveforms. We show that it is possible to hide sensitive information such as gender by generating new data, trained adversarially to maintain utility and realism.

Citations

Related