2021/07/05 by Tamás Gábor Csapó, Csapó, Tamás Gábor, László S. Tóth +5
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2107.02003
openalex publication_date 2021/07/05 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Articulatory information has been shown to be effective in improving the\nperformance of HMM-based and DNN-based text-to-speech synthesis. Speech\nsynthesis research focuses traditionally on text-to-speech conversion, when the\ninput is text or an estimated linguistic representation, and the target is\nsynthesized speech. However, a research field that has risen in the last decade\nis articulation-to-speech synthesis (with a target application of a Silent\nSpeech Interface, SSI), when the goal is to synthesize speech from some\nrepresentation of the movement of the articulatory organs. In this paper, we\nextend traditional (vocoder-based) DNN-TTS with articulatory input, estimated\nfrom ultrasound tongue images. We compare text-only, ultrasound-only, and\ncombined inputs. Using data from eight speakers, we show that that the combined\ntext and articulatory input can have advantages in limited-data scenarios,\nnamely, it may increase the naturalness of synthesized speech compared to\nsingle text input. Besides, we analyze the ultrasound tongue recordings of\nseveral speakers, and show that misalignments in the ultrasound transducer\npositioning can have a negative effect on the final synthesis performance.\n