vix.ing · top · new · best · stats

A Waveform Representation Framework for High-quality Statistical Parametric Speech Synthesis

2015/10/06 by Bo Fan, Fan, Bo, Siu Wa Lee +7
Computer Science · #68T10 #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.LG #cs.SD #msc:68T10

paper · pdf · doi:10.48550/arxiv.1510.01443

accepted and will appear in APSIPA2015; keywords: speech synthesis, LSTM-RNN, vocoder, phase, waveform, modeling

arxiv created 2015/10/06 · openalex publication_date 2015/10/06 · arxiv updated 2015/10/08 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years. Although perceptual studies have shown that phase spectrum is essential to the quality of synthesized speech, it is often ignored by using a minimum phase filter during synthesis and the speech quality suffers. To bypass this bottleneck in vocoded speech, this paper proposes a phase-embedded waveform representation framework and establishes a magnitude-phase joint modeling platform for high-quality SPSS. Our experiments on waveform reconstruction show that the performance is better than that of the widely-used STRAIGHT. Furthermore, the proposed modeling and synthesis platform outperforms a leading-edge, vocoded, deep bidirectional long short-term memory recurrent neural network (DBLSTM-RNN)-based baseline system in various objective evaluation metrics conducted.

Related