Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
2024/06/04 by Philip Anastassiou, Jiawei Chen, Anastassiou, Philip +89 · 164 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2406.02430
openalex publication_date 2024/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-TTS serves as a foundation model for speech generation and excels in speech in-context learning, achieving performance in speaker similarity and naturalness that matches ground truth human speech in both objective and subjective evaluations. With fine-tuning, we achieve even higher subjective scores across these metrics. Seed-TTS offers superior controllability over various speech attributes such as emotion and is capable of generating highly expressive and diverse speech for speakers in the wild. Furthermore, we propose a self-distillation method for speech factorization, as well as a reinforcement learning approach to enhance model robustness, speaker similarity, and controllability. We additionally present a non-autoregressive (NAR) variant of the Seed-TTS model, named Seed-TTSDiT, which utilizes a fully diffusion-based architecture. Unlike previous NAR-based TTS systems, Seed-TTSDiT does not depend on pre-estimated phoneme durations and performs speech generation through end-to-end processing. We demonstrate that this variant achieves comparable performance to the language model-based variant and showcase its effectiveness in speech editing. We encourage readers to listen to demos at \urlhttps://bytedancespeech.github.io/seedttstechreport.
Cited by
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
- Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization
- QuarkAudio Technical Report
- JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
- Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
- Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability
- GLM-TTS Technical Report
- Towards Interactive Intelligence for Digital Humans
- DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Qwen3.5-Omni Technical Report
- M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
- RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
- InstructAudio: Unified speech and music generation with natural language instruction
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
- Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- Step-Audio-EditX Technical Report
- SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
- UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
- LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
- Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
- MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
- IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
- CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
- Position: Towards Responsible Evaluation for Text-to-Speech
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
- MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
- HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- VoiceBridge: Designing Latent Bridge Models for General Speech Restoration at Scale
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
- Selective Classifier-free Guidance for Zero-shot Text-to-speech
- Qwen-Audio-3.0-Gen-Preview Technical Report
- Direct Preference Optimization for Speech Autoregressive Diffusion Models
- Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
- Group Relative Policy Optimization for Text-to-Speech with Large Language Models
- Explore the Reinforcement Learning for the LLM based ASR and TTS system
- Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
- Qwen3-Omni Technical Report
- Bridging the gap between training and inference in LM-based TTS models
- MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
- SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
- MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
- DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
- Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
- Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
- Audio Deepfake Verification
- VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- LibriQuote: A Speech Dataset of Fictional Character Utterances for Expressive Zero-Shot Speech Synthesis
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- VibeVoice Technical Report
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
- Long-Context Speech Synthesis with Context-Aware Memory
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
- MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
- M3PDB: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
- Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
- Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
- Adaptive Duration Model for Text Speech Alignment
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
- HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
- DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
- ERNIE 5.0 Technical Report
- Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
- Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
- TTS-1 Technical Report
- SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
- Step-Audio 2 Technical Report
- Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
- ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
- Differentiable Reward Optimization for LLM based TTS system
- Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
- Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
- InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
- RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
- MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
- EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
- Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
- Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
- VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
- Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
- UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
- VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
- A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
- SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
- DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
- Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
- UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
- DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
- MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
- Voice Cloning: Comprehensive Survey
- dots.tts Technical Report
- UniVocal: Unified Speech-Singing Code-Switching Synthesis
- From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
- Towards Flow-Matching-based TTS without Classifier-Free Guidance
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a 50K Budget
- Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
- DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
- Best-of-N TTS Evaluation is Confounded by ASR Family Alignment
- GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
- USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
- Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis
- P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Related