vix.ing · top · new · best · stats · spec

Unsupervised Cross-Domain Speech-to-Speech Conversion with\n Time-Frequency Consistency

2020/05/15 by Fabien Cardinaux, Khan, Mohammad Asif, Stefan Uhlich +7
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2005.07810

openalex publication_date 2020/05/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In recent years generative adversarial network (GAN) based models have been\nsuccessfully applied for unsupervised speech-to-speech conversion.The rich\ncompact harmonic view of the magnitude spectrogram is considered a suitable\nchoice for training these models with audio data. To reconstruct the speech\nsignal first a magnitude spectrogram is generated by the neural network, which\nis then utilized by methods like the Griffin-Lim algorithm to reconstruct a\nphase spectrogram. This procedure bears the problem that the generated\nmagnitude spectrogram may not be consistent, which is required for finding a\nphase such that the full spectrogram has a natural-sounding speech waveform. In\nthis work, we approach this problem by proposing a condition encouraging\nspectrogram consistency during the adversarial training procedure. We\ndemonstrate our approach on the task of translating the voice of a male speaker\nto that of a female speaker, and vice versa. Our experimental results on the\nLibrispeech corpus show that the model trained with the TF consistency provides\na perceptually better quality of speech-to-speech conversion.\n

Citations

Related