VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 67 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2406.05370
Abstract
This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity for the first time. Based on its predecessor, VALL-E, the new iteration introduces two significant enhancements: Repetition Aware Sampling refines the original nucleus sampling process by accounting for token repetition in the decoding history. It not only stabilizes the decoding but also circumvents the infinite loop issue. Grouped Code Modeling organizes codec codes into groups to effectively shorten the sequence length, which not only boosts inference speed but also addresses the challenges of long sequence modeling. Our experiments on the LibriSpeech and VCTK datasets show that VALL-E 2 surpasses previous systems in speech robustness, naturalness, and speaker similarity. It is the first of its kind to reach human parity on these benchmarks. Moreover, VALL-E 2 consistently synthesizes high-quality speech, even for sentences that are traditionally challenging due to their complexity or repetitive phrases. The advantages of this work could contribute to valuable endeavors, such as generating speech for individuals with aphasia or people with amyotrophic lateral sclerosis. See https://aka.ms/valle2 for demos of VALL-E 2.
Cited by
- Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
- SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
- STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
- VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
- Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
- Step-Audio-EditX Technical Report
- SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
- Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
- Position: Towards Responsible Evaluation for Text-to-Speech
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
- Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
- CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
- Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
- Teffic-Audio: Tell Fact from Fiction
- MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
- DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
- No Encore: Unlearning as Opt-Out in Music Generation
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- XMUspeech Systems for the ASVspoof 5 Challenge
- DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
- M3PDB: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
- MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
- QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
- Scalable Controllable Accented TTS
- Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
- Next Tokens Denoising for Speech Synthesis
- DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
- Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
- Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
- DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
- Speaking images. A novel framework for the automated self-description of artworks
- Voice Adaptation for Swiss German
- VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
- Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
- DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
- OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
- Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech
- DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
- AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
- Voice Cloning: Comprehensive Survey
- MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals
- AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
- Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a 50K Budget
- GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
Related