2021/02/19 by Brooke Stephenson, Thomas Hueber, Stephenson, Brooke +5
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2102.09914
openalex publication_date 2021/02/19 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
The prosody of a spoken word is determined by its surrounding context. In\nincremental text-to-speech synthesis, where the synthesizer produces an output\nbefore it has access to the complete input, the full context is often unknown\nwhich can result in a loss of naturalness in the synthesized speech. In this\npaper, we investigate whether the use of predicted future text can attenuate\nthis loss. We compare several test conditions of next future word: (a) unknown\n(zero-word), (b) language model predicted, (c) randomly predicted and (d)\nground-truth. We measure the prosodic features (pitch, energy and duration) and\nfind that predicted text provides significant improvements over a zero-word\nlookahead, but only slight gains over random-word lookahead. We confirm these\nresults with a perceptive test.\n