VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
2025/05/26 by Peng, Puyuan, Li, Shang-Wen, Mohamed, Abdelrahman +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2505.19462
Abstract
We present VoiceStar, the first zero-shot TTS model that achieves both output duration control and extrapolation. VoiceStar is an autoregressive encoder-decoder neural codec language model, that leverages a novel Progress-Monitoring Rotary Position Embedding (PM-RoPE) and is trained with Continuation-Prompt Mixed (CPM) training. PM-RoPE enables the model to better align text and speech tokens, indicates the target duration for the generated speech, and also allows the model to generate speech waveforms much longer in duration than those seen during. CPM training also helps to mitigate the training/inference mismatch, and significantly improves the quality of the generated speech in terms of speaker similarity and intelligibility. VoiceStar outperforms or is on par with current state-of-the-art models on short-form benchmarks such as Librispeech and Seed-TTS, and significantly outperforms these models on long-form/extrapolation benchmarks (20-50s) in terms of intelligibility and naturalness. Code and models: https://github.com/jasonppy/VoiceStar. Audio samples: https://jasonppy.github.io/VoiceStarweb
Citations
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
- Scaling Rich Style-Prompted Text-to-Speech Datasets
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
- Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
- Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
- Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
- Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
- DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
- IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
- HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
- Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
- Enabling Real-Time Conversations with Minimal Training Costs
- Moshi: a speech-text foundation model for real-time dialogue
- SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
- Language Model Can Listen While Speaking
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
- E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
- Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
- VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
- Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
- Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
- VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
- BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
- Natural language guidance of high-fidelity text-to-speech with synthetic annotations
- SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
- ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
- Audiobox: Unified Audio Generation with Natural Language Prompts
- Zipformer: A faster and better encoder for automatic speech recognition
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
- Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
- PromptTTS 2: Describing and Generating Voices with Text Prompt
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
- Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
- Simple and Controllable Music Generation
- PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions
- VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
- SoundStorm: Efficient Parallel Audio Generation
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
- Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
- InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Scalable Diffusion Models with Transformers
- Robust Speech Recognition via Large-Scale Weak Supervision
- PromptTTS: Controllable Text-to-Speech with Text Descriptions
- High Fidelity Neural Audio Compression
- AudioGen: Textually Guided Audio Generation
- AudioLM: a Language Modeling Approach to Audio Generation
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- SoundStream: An End-to-End Neural Audio Codec
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Generative Spoken Language Modeling from Raw Audio
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Attention Is All You Need
- Neural Machine Translation by Jointly Learning to Align and Translate
Cited by
Related