FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
2020/06/08 by Yi Ren, Chenxu Hu, Ren, Yi +12 · 154 citations
Computer Science · Engineering · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG #cs.SD #eess.AS
paper · pdf · doi:10.48550/arxiv.2006.04558
Accepted by ICLR 2021
arxiv created 2022/08/08 · arxiv updated 2022/08/09
Abstract
Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher model for duration prediction (to provide more information as input) and knowledge distillation (to simplify the data distribution in output), which can ease the one-to-many mapping problem (i.e., multiple speech variations correspond to the same text) in TTS. However, FastSpeech has several disadvantages: 1) the teacher-student distillation pipeline is complicated and time-consuming, 2) the duration extracted from the teacher model is not accurate enough, and the target mel-spectrograms distilled from teacher model suffer from information loss due to data simplification, both of which limit the voice quality. In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the simplified output from teacher, and 2) introducing more variation information of speech (e.g., pitch, energy and more accurate duration) as conditional inputs. Specifically, we extract duration, pitch and energy from speech waveform and directly take them as conditional inputs in training and use predicted values in inference. We further design FastSpeech 2s, which is the first attempt to directly generate speech waveform from text in parallel, enjoying the benefit of fully end-to-end inference. Experimental results show that 1) FastSpeech 2 achieves a 3x training speed-up over FastSpeech, and FastSpeech 2s enjoys even faster inference speed; 2) FastSpeech 2 and 2s outperform FastSpeech in voice quality, and FastSpeech 2 can even surpass autoregressive models. Audio samples are available at https://speechresearch.github.io/fastspeech2/.
Citations
Cited by
- Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
- The Closer, The Better: How does voice similarity affect our preference for synthetic voices?
- InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
- Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
- Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability
- DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
- Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
- M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
- STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
- VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
- CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
- Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
- SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
- The Impact of Prosodic Segmentation on Speech Synthesis of Spontaneous Speech
- Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated Approach
- Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- Review of end-to-end speech synthesis technology based on deep learning
- Bayesian Speech synthesizers Can Learn from Multiple Teachers
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
- ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
- Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
- Universal Discrete-Domain Speech Enhancement
- Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- Position: Towards Responsible Evaluation for Text-to-Speech
- A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance
- STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
- LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
- UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
- Metamorphic Testing for Audio Content Moderation Software
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- Emotional Text-To-Speech in Japanese Using Artificially Augmented Dataset
- AUDDT: Audio Unified Deepfake Detection Benchmark Toolkit
- PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
- OLaPh: Optimal Language Phonemizer
- SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
- VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
- TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
- Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
- Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
- RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
- Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
- Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
- MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions
- ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
- XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
- Improved Dysarthric Speech to Text Conversion via TTS Personalization
- DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
- UniTalker: Conversational Speech-Visual Synthesis
- Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
- Inference-time Scaling for Diffusion-based Audio Super-resolution
- Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
- Adaptive Duration Model for Text Speech Alignment
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
- Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
- Step-Audio 2 Technical Report
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
- EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
- Supporting SENĆOTEN Language Documentation Efforts with Automatic Speech Recognition
- Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
- LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
- Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
- Eigenvoice Synthesis based on Model Editing for Speaker Generation
- First Steps Towards Voice Anonymization for Code-Switching Speech
- Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
- Multi-interaction TTS toward professional recording reproduction
- Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
- StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
- JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
- Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
- SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
- LeVo: High-Quality Song Generation with Multi-Preference Alignment
- Selecting N-lowest scores for training MOS prediction models
- Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
- RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
- DeepSinger: Singing Voice Synthesis with Data Mined From the Web
- VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
- InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
- TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
- Instance-Specific Test-Time Training for Speech Editing in the Wild
- S2ST-Omni: An Efficient Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Progressive Fine-tuning
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
- LightSpeech: Lightweight and Fast Text to Speech with Neural Architecture Search
- Voice Impression Control in Zero-Shot TTS
- Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
- LLAMAPIE: Proactive In-Ear Conversation Assistants
- Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
- Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
- DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
- Length Aware Speech Translation for Video Dubbing
- Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
- RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
- VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
- Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
- Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
- A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
- DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
- SpeakStream: Streaming Text-to-Speech with Interleaved Data
- CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
- Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
- MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
- Private kNN-VC: Interpretable Anonymization of Converted Speech
- LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
- Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
- Pairwise Evaluation of Accent Similarity in Speech Synthesis
- Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
- Score-Based Training for Energy-Based TTS Models
- Fast, Not Fancy: Rethinking G2P with Rich Data and Rule-Based Models
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
- Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
- On the Cost and Benefits of Training Context with Utterance or Full Conversation Training: A Comparative Stud
- Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations
- AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
- Voice Cloning: Comprehensive Survey
- Perceptual Implications of Automatic Anonymization in Pathological Speech
- AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
- Versatile Framework for Song Generation with Prompt-based Control
- Using Phonemes in cascaded S2S translation pipeline
- DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
- Generalized Audio Deepfake Detection Using Frame-level Latent Information Entropy
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
- AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
- Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
- Cellular Development Follows the Path of Minimum Action
- Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis
- SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
- P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Related