vix.ing · top · new · best · stats · spec

Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen\n Speaker and Recording Conditions

2020/08/09 by Dipjyoti Paul, Paul, Dipjyoti, Yannis Pantazis +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2008.05289

openalex publication_date 2020/08/09 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

Recent advancements in deep learning led to human-level performance in\nsingle-speaker speech synthesis. However, there are still limitations in terms\nof speech quality when generalizing those systems into multiple-speaker models\nespecially for unseen speakers and unseen recording qualities. For instance,\nconventional neural vocoders are adjusted to the training speaker and have poor\ngeneralization capabilities to unseen speakers. In this work, we propose a\nvariant of WaveRNN, referred to as speaker conditional WaveRNN (SC-WaveRNN). We\ntarget towards the development of an efficient universal vocoder even for\nunseen speakers and recording conditions. In contrast to standard WaveRNN,\nSC-WaveRNN exploits additional information given in the form of speaker\nembeddings. Using publicly-available data for training, SC-WaveRNN achieves\nsignificantly better performance over baseline WaveRNN on both subjective and\nobjective metrics. In MOS, SC-WaveRNN achieves an improvement of about 23% for\nseen speaker and seen recording condition and up to 95% for unseen speaker and\nunseen condition. Finally, we extend our work by implementing a multi-speaker\ntext-to-speech (TTS) synthesis similar to zero-shot speaker adaptation. In\nterms of performance, our system has been preferred over the baseline TTS\nsystem by 60% over 15.5% and by 60.9% over 32.6%, for seen and unseen speakers,\nrespectively.\n

Related