Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
2021/05/13 by Vadim Popov, Popov Va, Ivan Vovk +8 · 89 citations
Computer Science · Mathematics · Physics and Astronomy · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Opinion Dynamics and Social Influence #Speech and Audio Processing #cs.CL #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2105.06337
openalex publication_date 2021/05/13 · openalex created_date 2021/05/24 · arxiv created 2021/08/05 · arxiv updated 2021/08/06 · openalex updated_date 2026/07/28
Abstract
Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these techniques allowing for flexible inference schemes. In this paper we introduce Grad-TTS, a novel text-to-speech model with score-based decoder producing mel-spectrograms by gradually transforming noise predicted by encoder and aligned with text input by means of Monotonic Alignment Search. The framework of stochastic differential equations helps us to generalize conventional diffusion probabilistic models to the case of reconstructing data from noise with different parameters and allows to make this reconstruction flexible by explicitly controlling trade-off between sound quality and inference speed. Subjective human evaluation shows that Grad-TTS is competitive with state-of-the-art text-to-speech approaches in terms of Mean Opinion Score. We will make the code publicly available shortly.
Citations
Cited by
- Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
- CoPHo: Classifier-guided Conditional Topology Generation with Persistent Homology
- Adapting Speech Language Model to Singing Voice Synthesis
- Personalized Federated Distillation Assisted Vehicle Edge Caching Strategy
- Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
- NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation
- Solving Diffusion Inverse Problems with Restart Posterior Sampling
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
- Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
- Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Adaptive Discretization for Consistency Models
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- OO-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
- A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and Masking
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
- Environment-Aware Satellite Image Generation with Diffusion Models
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
- SimDiff: Simulator-constrained Diffusion Model for Physically Plausible Motion Generation
- Training Flow Matching Models with Reliable Labels via Self-Purification
- Discrete-Time Diffusion-Like Models for Speech Synthesis
- An Octave-based Multi-Resolution CQT Architecture for Diffusion-based Audio Generation
- Real-Time Streaming Mel Vocoding with Generative Flow Matching
- Defending Diffusion Models Against Membership Inference Attacks via Higher-Order Langevin Dynamics
- MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
- WildSpoof Challenge Evaluation Plan
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- STADI: Fine-Grained Step-Patch Diffusion Parallelism for Heterogeneous GPUs
- A Survey on Neural Speech Synthesis
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- Audio-Guided Visual Editing with Complex Multi-Modal Prompts
- Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System
- DiffIER: Optimizing Diffusion Models with Iterative Error Reduction
- Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
- Transient Noise Removal via Diffusion-based Speech Inpainting
- UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
- Improved Dysarthric Speech to Text Conversion via TTS Personalization
- Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
- Flow Matching Policy Gradients
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
- EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
- How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
- Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
- Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
- LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
- PresentAgent: Multimodal Agent for Presentation Video Generation
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
- You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
- SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
- Selecting N-lowest scores for training MOS prediction models
- Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
- JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
- RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
- VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
- TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- Federated Learning Assisted Edge Caching Scheme Based on Lightweight Architecture DDPM
- Denoising Diffusion Gamma Models
- SegDiff: Image Segmentation with Diffusion Probabilistic Models
- Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages
- InfiniteAudio: Infinite-Length Audio Generation with Consistency
- ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
- ZeroSep: Separate Anything in Audio with Zero Training
- Almost Linear Convergence under Minimal Score Assumptions: Quantized Transition Diffusion
- ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
- Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer
- CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
- MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
- Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
- Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
- Language translation, and change of accent for speech-to-speech task using diffusion model
- Voice Cloning: Comprehensive Survey
- Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
- DRAGON: Distributional Rewards Optimize Diffusion Generative Models
- Generalized Audio Deepfake Detection Using Frame-level Latent Information Entropy
- SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
- On the Design of Diffusion-based Neural Speech Codecs
- P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Related