2019/04/08 by Lauri Juvela, Bajibabu Bollepalli, Juvela, Lauri +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1904.03976
openalex publication_date 2019/04/08 · openalex created_date 2022/07/24 · openalex updated_date 2026/07/28
Recent advances in neural network -based text-to-speech have reached human\nlevel naturalness in synthetic speech. The present sequence-to-sequence models\ncan directly map text to mel-spectrogram acoustic features, which are\nconvenient for modeling, but present additional challenges for vocoding (i.e.,\nwaveform generation from the acoustic features). High-quality synthesis can be\nachieved with neural vocoders, such as WaveNet, but such autoregressive models\nsuffer from slow sequential inference. Meanwhile, their existing parallel\ninference counterparts are difficult to train and require increasingly large\nmodel sizes. In this paper, we propose an alternative training strategy for a\nparallel neural vocoder utilizing generative adversarial networks, and\nintegrate a linear predictive synthesis filter into the model. Results show\nthat the proposed model achieves significant improvement in inference speed,\nwhile outperforming a WaveNet in copy-synthesis quality.\n