Tacotron: Towards End-to-End Speech Synthesis
2017/03/29 by Yuxuan Wang, Wang, Yuxuan, RJ Skerry-Ryan +26 · 2 voices · 115 citations
Computer Science · #cs.CL #cs.LG #cs.SD
paper · pdf · doi:10.48550/arxiv.1703.10135
Submitted to Interspeech 2017. v2 changed paper title to be consistent with our conference submission (no content change other than typo fixes)
arxiv created 2017/04/06 · arxiv updated 2017/04/10
Abstract
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain brittle design choices. In this paper, we present Tacotron, an end-to-end generative text-to-speech model that synthesizes speech directly from characters. Given <text, audio> pairs, the model can be trained completely from scratch with random initialization. We present several key techniques to make the sequence-to-sequence framework perform well for this challenging task. Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English, outperforming a production parametric system in terms of naturalness. In addition, since Tacotron generates speech at the frame level, it's substantially faster than sample-level autoregressive methods.
Citations
Cited by
- Towards High-Level Semantic Intelligence
- The Closer, The Better: How does voice similarity affect our preference for synthetic voices?
- Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- GLM-TTS Technical Report
- DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
- All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
- Speech-to-Singing Conversion in an Encoder-Decoder Framework
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Is Phase Really Needed for Weakly-Supervised Dereverberation ?
- Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
- Statistical Parametric Speech Synthesis Using Generative Adversarial Networks Under A Multi-task Learning Framework
- Adversarially Trained Autoencoders for Parallel-Data-Free Voice Conversion
- Unified Mandarin TTS Front-end Based on Distilled BERT Model
- Listening while Speaking: Speech Chain by Deep Learning
- Bayesian Speech synthesizers Can Learn from Multiple Teachers
- emg2speech: synthesizing speech from electromyography using self-supervised speech models
- Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- Position: Towards Responsible Evaluation for Text-to-Speech
- Metamorphic Testing for Audio Content Moderation Software
- ArFake: A Robust Framework for Multi-Dialect Arabic Speech Spoofing Detection Benchmark
- Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features
- AUDDT: Audio Unified Deepfake Detection Benchmark Toolkit
- Whispered and Lombard Neural Speech Synthesis
- FastSpeech: Fast, Robust and Controllable Text to Speech
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- Attention Forcing for Sequence-to-sequence Model Training
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems
- MelNet: A Generative Model for Audio in the Frequency Domain
- AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
- Phoneme-based Distribution Regularization for Speech Enhancement
- Emotional Voice Conversion using Multitask Learning with Text-to-speech
- FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks
- FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
- Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis
- Rhythm-Flexible Voice Conversion without Parallel Data Using Cycle-GAN over Phoneme Posteriorgram Sequences
- FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
- Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
- Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial Training
- Text to Speech System for Meitei Mayek Script
- Machine Speech Chain with One-shot Speaker Adaptation
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- Variational Bi-LSTMs
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
- Learning Neural Vocoder from Range-Null Space Decomposition
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- A unified sequence-to-sequence front-end model for Mandarin text-to-speech synthesis
- Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems
- Large-scale Speaker Retrieval on Random Speaker Variability Subspace
- PortaSpeech: Portable and High-Quality Generative Text-to-Speech
- Mathematical Vocoder Algorithm : Modified Spectral Inversion for Efficient Neural Speech Synthesis
- Towards Robust Neural Vocoding for Speech Generation: A Survey
- Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
- Noise Adaptive Speech Enhancement using Domain Adversarial Training
- Step-Audio 2 Technical Report
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
- Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet
- Recent Advances and Trends in Multimodal Deep Learning: A Review
- TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
- Facetron: A Multi-speaker Face-to-Speech Model based on Cross-modal Latent Representations
- Semi-Supervised Neural Architecture Search
- Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
- JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
- Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations
- IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection
- Speaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention
- V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
- RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
- Fitting New Speakers Based on a Short Untranscribed Sample
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
- Learning neural trans-dimensional random field language models with noise-contrastive estimation
- SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
- LSTM Acoustic Models Learn to Align and Pronounce with Graphemes
- End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator
- Multi-Reference Neural TTS Stylization with Adversarial Cycle Consistency
- Transfer Learning from Monolingual ASR to Transcription-free Cross-lingual Voice Conversion
- Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
- Embodied Self-supervised Learning by Coordinated Sampling and Training
- SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
- Tell me Habibi, is it Real or Fake?
- Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
- GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
- SpeakStream: Streaming Text-to-Speech with Interleaved Data
- Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
- Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
- Action2Dialogue: Generating Character-Centric Narratives from Scene-Level Prompts
- Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Building Multi lingual TTS using Cross Lingual Voice Conversion
- Effective parameter estimation methods for an ExcitNet model in generative text-to-speech systems
- On the Cost and Benefits of Training Context with Utterance or Full Conversation Training: A Comparative Stud
- Beyond Identity: A Generalizable Approach for Deepfake Audio Detection
- Voice Cloning: Comprehensive Survey
- AdaSpeech 2: Adaptive Text to Speech with Untranscribed Data
- AI based Presentation Creator With Customized Audio Content Delivery
- Script2Screen: Supporting Dialogue Scriptwriting with Interactive Audiovisual Generation
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
- AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
- Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
- P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Discussions
Related