All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
2025/12/12 by Moriya, Takafumi, Mimura, Masato, Tanaka, Tomohiro +3
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2512.11543
Abstract
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy.
Citations
- Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers
- Alignment-Free Training for Transducer-based Multi-Talker ASR
- Theory, Analysis, and Best Practices for Sigmoid Self-Attention
- Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
- Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
- Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
- XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
- Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
- Mamba in Speech: Towards an Alternative to Self-Attention
- CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
- End-to-End Joint Target and Non-Target Speakers ASR
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers
- Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
- Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR
- End-to-End Speech Recognition: A Survey
- Fast and accurate factorized neural transducer for text adaption of end-to-end speech recognition models
- Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers
- Monotonic segmental attention for automatic speech recognition
- Streaming Target-Speaker ASR with Neural Transducer
- Pruned RNN-T for fast, memory-efficient ASR training
- An Investigation of Monotonic Transducers for Large-Scale Automatic Speech Recognition
- Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
- Streaming Multi-Talker ASR with Token-Level Serialized Output Training
- A Study of Transducer based End-to-End ASR with ESPnet: Architecture, Auxiliary Loss and Decoding Strategies
- Recent Advances in End-to-End Automatic Speech Recognition
- Speech Summarization using Restricted Self-Attention
- Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution
- Advancing RNN Transducer Technology for Speech Recognition
- Alignment Knowledge Distillation for Online Streaming Attention-based Speech Recognition
- A Better and Faster End-to-End Model for Streaming ASR
- Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Serialized Output Training for End-to-End Overlapped Speech Recognition
- Deliberation Model Based Two-Pass End-to-End Speech Recognition
- Hybrid Autoregressive Transducer (hat)
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- Streaming automatic speech recognition with the transformer model
- SpecAugment on Large Scale Datasets
- A Density Ratio Approach to Language Model Fusion in End-To-End\n Automatic Speech Recognition
- A comparison of end-to-end models for long-form speech recognition
- Root Mean Square Layer Normalization
- Two-Pass End-to-End Speech Recognition
- ESPnet: End-to-End Speech Processing Toolkit
- Monotonic Chunkwise Attention
- An analysis of incorporating an external language model into a sequence-to-sequence model
- Attention Is All You Need
- Tacotron: Towards End-to-End Speech Synthesis
- Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
- Layer Normalization
- On Multiplicative Integration with Recurrent Neural Networks
- Highway Long Short-Term Memory RNNs for Distant Speech Recognition
- Neural Machine Translation of Rare Words with Subword Units
- On Using Monolingual Corpora in Neural Machine Translation
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Adam: A Method for Stochastic Optimization
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Sequence Transduction with Recurrent Neural Networks
Related