2019/07/01 by Ahmed Mustafa, Mustafa, Ahmed, Arijit Biswas +7
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1907.00772
openalex publication_date 2019/07/01 · openalex created_date 2022/07/22 · openalex updated_date 2026/07/28
Classical parametric speech coding techniques provide a compact\nrepresentation for speech signals. This affords a very low transmission rate\nbut with a reduced perceptual quality of the reconstructed signals. Recently,\nautoregressive deep generative models such as WaveNet and SampleRNN have been\nused as speech vocoders to scale up the perceptual quality of the reconstructed\nsignals without increasing the coding rate. However, such models suffer from a\nvery slow signal generation mechanism due to their sample-by-sample modelling\napproach. In this work, we introduce a new methodology for neural speech\nvocoding based on generative adversarial networks (GANs). A fake speech signal\nis generated from a very compressed representation of the glottal excitation\nusing conditional GANs as a deep generative model. This fake speech is then\nrefined using the LPC parameters of the original speech signal to obtain a\nnatural reconstruction. The reconstructed speech waveforms based on this\napproach show a higher perceptual quality than the classical vocoder\ncounterparts according to subjective and objective evaluation scores for a\ndataset of 30 male and female speakers. Moreover, the usage of GANs enables to\ngenerate signals in one-shot compared to autoregressive generative models. This\nmakes GANs promising for exploration to implement high-quality neural vocoders.\n