vix.ing · top · new · best · stats · spec

Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection

2024/08/30 by İsmail Rasim Ülgen, Shreeram Suresh Chandra, Ulgen, Ismail Rasim +5 · 1 citation
Biochemistry, Genetics and Molecular Biology · Engineering · #Audio and Speech Processing (eess.AS) #DNA and Biological Computing #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Modular Robots and Swarm Intelligence #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2408.17432

openalex publication_date 2024/08/30 · openalex created_date 2024/10/19 · openalex updated_date 2026/07/28

Abstract

Synthesizing the voices of unseen speakers remains a persisting challenge in multi-speaker text-to-speech (TTS). Existing methods model speaker characteristics through speaker conditioning during training, leading to increased model complexity and limiting reproducibility and accessibility. A low-complexity alternative would broaden the reach of speech synthesis research, particularly in settings with limited computational and data resources. To this end, we propose SelectTTS, a simple and effective alternative. SelectTTS selects appropriate frames from the target speaker and decodes them using frame-level self-supervised learning (SSL) features. We demonstrate that this approach can effectively capture speaker characteristics for unseen speakers and achieves performance comparable to state-of-the-art multi-speaker TTS frameworks on both objective and subjective metrics. By directly selecting frames from the target speaker's speech, SelectTTS enables generalization to unseen speakers with significantly lower model complexity. Experimental results show that the proposed approach achieves performance comparable to state-of-the-art systems such as XTTS-v2 and VALL-E, while requiring over 8x fewer parameters and 270x less training data. Moreover, it demonstrates that frame selection with SSL features offers an efficient path to low-complexity, high-quality multi-speaker TTS.

Cited by

Related