vix.ing · top · new · best · stats

Speaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention

2018/10/29 by Bajibabu Bollepalli, Bollepalli, Bajibabu, Lauri Juvela +3 · 2 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1810.12051

5 pages, 5 figures. Submitted to ICASSP 2019

arxiv created 2018/10/29 · openalex publication_date 2018/10/29 · arxiv updated 2018/10/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration properties directly from text and speech. Since these models are entirely data-driven, they need large amounts of data to generate synthetic speech with good quality. However, in challenging speaking styles, such as Lombard speech, it is difficult to record sufficiently large speech corpora. Therefore, in this study we propose a transfer learning method to adapt a sequence-to-sequence based TTS system of normal speaking style to Lombard style. Moreover, we experiment with a WaveNet vocoder in synthesis of Lombard speech. We conducted subjective evaluations to assess the performance of the adapted TTS systems. The subjective evaluation results indicated that an adaptation system with the WaveNet vocoder clearly outperformed the conventional deep neural network based TTS system in synthesis of Lombard speech.

Citations

Cited by

Related