vix.ing · top · new · best · stats · spec

Generative Adversarial Network based Speaker Adaptation for High Fidelity WaveNet Vocoder

2018/12/06 by Qiao Tian, Tian, Qiao, Xucheng Wan +3
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1812.02339

5 pages, 4 figure, 1 table, 6 equations

openalex publication_date 2018/12/06 · openalex created_date 2018/12/11 · arxiv created 2019/07/19 · arxiv updated 2019/07/22 · openalex updated_date 2026/07/28

Abstract

Although state-of-the-art parallel WaveNet has addressed the issue of real-time waveform generation, there remains problems. Firstly, due to the noisy input signal of the model, there is still a gap between the quality of generated and natural waveforms. Secondly, a parallel WaveNet is trained under a distillation framework, which makes it tedious to adapt a well trained model to a new speaker. To address these two problems, in this paper we propose an end-to-end adaptation method based on the generative adversarial network (GAN), which can reduce the computational cost for the training of new speaker adaptation. Our subjective experiments shows that the proposed training method can further reduce the quality gap between generated and natural waveforms.

Citations

Related