vix.ing · top · new · best · stats · spec

Controllable Prosody Generation With Partial Inputs

2023/03/14 by Dan Andrei Iliescu, Iliescu, Dan Andrei, Devang Savita Ram Mohan +5
Computer Science · Neuroscience · #AI in cancer detection #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Brain Tumor Detection and Classification #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Medical Image Segmentation Techniques #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2303.09446

openalex publication_date 2023/03/14 · openalex created_date 2023/03/19 · openalex updated_date 2026/07/28

Abstract

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can modify the output quickly and precisely. To solve this, we introduce a novel framework whereby the user provides partial inputs and the generative model generates the missing features. We propose a model that is specifically designed to encode partial prosodic features and output complete audio. We show empirically that our model displays two essential qualities of a human-in-the-loop control mechanism: efficiency and robustness. With even a very small number of input values (~4), our model enables users to improve the quality of the output significantly in terms of listener preference (4:1).

Related