vix.ing · top · new · best · stats · spec

Synthesis of Tongue Motion and Acoustics From Text Using a Multimodal Articulatory Database

2016/12/31 by Ingmar Steiner, Sébastien Le Maguer, Sebastien Le Maguer +1
Computer Science · Medicine · Psychology · #Articulation (sociology) #Articulatory phonetics #Modality (human–computer interaction) #Motion (physics) #Motion capture #Parametric statistics #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech synthesis #Tongue #Voice and Speech Disorders #cs.HC

paper · pdf · doi:10.1109/taslp.2017.2756818

published as IEEE/ACM Transactions on Audio, Speech, and Language Processing 25 (2017) 2351 - 2361

openalex created_date 2017/01/13 · openalex publication_date 2017/11/23 · arxiv created 2018/04/13 · arxiv updated 2018/04/17 · openalex updated_date 2026/08/05

Abstract

We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a three-dimensional model of the tongue surface to an articulatory dataset and training a statistical parametric speech synthesis system directly on the tongue model parameters. We evaluate the model at every step by comparing the spatial coordinates of predicted articulatory movements against the reference data. The results indicate a global mean Euclidean distance of less than 2.8 mm, and our approach can be adapted to add an articulatory modality to conventional TTS applications without the need for extra data.

Citations