2020/08/16 by Hyun-Wook Yoon, Yoon, Hyun-Wook, Lee, Sang-Hoon +4
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2008.06867
openalex publication_date 2020/08/16 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In recent works, a flow-based neural vocoder has shown significant\nimprovement in real-time speech generation task. The sequence of invertible\nflow operations allows the model to convert samples from simple distribution to\naudio samples. However, training a continuous density model on discrete audio\ndata can degrade model performance due to the topological difference between\nlatent and actual distribution. To resolve this problem, we propose audio\ndequantization methods in flow-based neural vocoder for high fidelity audio\ngeneration. Data dequantization is a well-known method in image generation but\nhas not yet been studied in the audio domain. For this reason, we implement\nvarious audio dequantization methods in flow-based neural vocoder and\ninvestigate the effect on the generated audio. We conduct various objective\nperformance assessments and subjective evaluation to show that audio\ndequantization can improve audio generation quality. From our experiments,\nusing audio dequantization produces waveform audio with better harmonic\nstructure and fewer digital artifacts.\n