Conformer: Convolution-augmented Transformer for Speech Recognition
2020/10/25 by Anmol Gulati, James Qin, Chung‐Cheng Chiu +8 · 143 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
paper · doi:10.21437/interspeech.2020-3015
openalex publication_date 2020/10/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).Transformer models are good at capturing content-based global interactions, while CNNs exploit local features effectively.In this work, we achieve the best of both worlds by studying how to combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence in a parameter-efficient way.To this regard, we propose the convolution-augmented transformer for speech recognition, named Conformer.Conformer significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies.On the widely used LibriSpeech benchmark, our model achieves WER of 2.1%/4.3%without using a language model and 1.9%/3.9%with an external language model on test/testother.We also observe competitive performance of 2.7%/6.3%with a small model of only 10M parameters.
Citations
Cited by
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection
- BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
- Bayesian Speech synthesizers Can Learn from Multiple Teachers
- Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
- V-SAT: Video Subtitle Annotation Tool
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- Data-Centric Lessons To Improve Speech-Language Pretraining
- VBx for End-to-End Neural and Clustering-based Diarization
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- Can large audio language models understand child stuttering speech? speech summarization, and source separation
- MLMA: Towards Multilingual ASR With Mamba-based Architectures
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
- Fighter: Unveiling the Graph Convolutional Nature of Transformers in Time Series Modeling
- Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
- Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
- LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
- Universal Discrete-Domain Speech Enhancement
- Target speaker anonymization in multi-speaker recordings
- Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
- Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
- Layer Pruning on Demand with Intermediate CTC
- Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
- XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
- Position: Towards Responsible Evaluation for Text-to-Speech
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems
- Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
- SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
- ASTROCO: Self-Supervised Conformer-Style Transformers for Light-Curve Embeddings
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
- LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
- ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
- InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
- Retrieval Augmented Generation based context discovery for ASR
- Teffic-Audio: Tell Fact from Fiction
- Cross-attention conformer for context modeling in speech enhancement for ASR
- A perceptual similarity space for speech based on self-supervised speech representations
- FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
- Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
- HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
- Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- BiLCNet : BiLSTM-Conformer Network for Encrypted Traffic Classification with 5G SA Physical Channel Records
- SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
- Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
- TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
- Interpreting the Role of Visemes in Audio-Visual Speech Recognition
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
- BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
- Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
- Pushing the Limits of End-to-End Diarization
- UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition
- FCPE: A Fast Context-based Pitch Estimation Model
- A long-form single-speaker real-time MRI speech dataset and benchmark
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
- GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
- SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- Error Analysis in a Modular Meeting Transcription System
- N-Singer: A Non-Autoregressive Korean Singing Voice Synthesis System for Pronunciation Enhancement
- Unified Learnable 2D Convolutional Feature Extraction for ASR
- Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR
- Graph Connectionist Temporal Classification for Phoneme Recognition
- XMUspeech Systems for the ASVspoof 5 Challenge
- Contextualized Token Discrimination for Speech Search Query Correction
- PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
- Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
- Human Motion Video Generation: A Survey
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
- NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
- Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
- AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
- From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
- Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
- CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
- Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
- Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
- Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
- HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- BConformeR: A Conformer Based on Mutual Sampling for Unified Prediction of Continuous and Discontinuous Antibody Binding Sites
- Pretrained Conformers for Audio Fingerprinting and Retrieval
- Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
- Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions
- A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models
- A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
- Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
- Joint decoding method for controllable contextual speech recognition based on Speech LLM
- Munsit at NADI 2025 Shared Task 2: Pushing the Boundaries of Multidialectal Arabic ASR with Weakly Supervised Pretraining and Continual Supervised Fine-tuning
- Revealing the Role of Audio Channels in ASR Performance Degradation
- DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
- Auditory Intelligence: Understanding the World Through Sound
- Score-Informed BiLSTM Correction for Refining MIDI Velocity in Automatic Piano Transcription
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
- SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
- A Study on Regularization-Based Continual Learning Methods for Indic ASR
- DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
- SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
- LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
- Adaptive Knowledge Distillation for Device-Directed Speech Detection
- CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
- Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
- Self-Improvement for Audio Large Language Model using Unlabeled Speech
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Delta Keyword Transformer: Bringing Transformers to the Edge through Dynamically Pruned Multi-Head Self-Attention
Related