2016/01/26 by Prasanna Kumar Muthukumar, Muthukumar, Prasanna Kumar, Alan W. Black +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
paper · pdf · doi:10.48550/arxiv.1601.07215
openalex publication_date 2016/01/26 · openalex created_date 2022/09/02 · openalex updated_date 2026/07/28
In the last two years, there have been numerous papers that have looked into\nusing Deep Neural Networks to replace the acoustic model in traditional\nstatistical parametric speech synthesis. However, far less attention has been\npaid to approaches like DNN-based postfiltering where DNNs work in conjunction\nwith traditional acoustic models. In this paper, we investigate the use of\nRecurrent Neural Networks as a potential postfilter for synthesis. We explore\nthe possibility of replacing existing postfilters, as well as highlight the\nease with which arbitrary new features can be added as input to the postfilter.\nWe also tried a novel approach of jointly training the Classification And\nRegression Tree and the postfilter, rather than the traditional approach of\ntraining them independently.\n