Conformer: Convolution-augmented Transformer for Speech Recognition
2020/10/25 by Anmol Gulati, James Qin, Chung‐Cheng Chiu +8 · 346 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
paper · doi:10.21437/interspeech.2020-3015
openalex publication_date 2020/10/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).Transformer models are good at capturing content-based global interactions, while CNNs exploit local features effectively.In this work, we achieve the best of both worlds by studying how to combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence in a parameter-efficient way.To this regard, we propose the convolution-augmented transformer for speech recognition, named Conformer.Conformer significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies.On the widely used LibriSpeech benchmark, our model achieves WER of 2.1%/4.3%without using a language model and 1.9%/3.9%with an external language model on test/testother.We also observe competitive performance of 2.7%/6.3%with a small model of only 10M parameters.
Citations
Cited by
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
- ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection
- BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
- Bayesian Speech synthesizers Can Learn from Multiple Teachers
- Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
- V-SAT: Video Subtitle Annotation Tool
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- Data-Centric Lessons To Improve Speech-Language Pretraining
- VBx for End-to-End Neural and Clustering-based Diarization
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- Can large audio language models understand child stuttering speech? speech summarization, and source separation
- MLMA: Towards Multilingual ASR With Mamba-based Architectures
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
- Fighter: Unveiling the Graph Convolutional Nature of Transformers in Time Series Modeling
- Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
- Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
- LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
- Universal Discrete-Domain Speech Enhancement
- Target speaker anonymization in multi-speaker recordings
- Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
- Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
- Layer Pruning on Demand with Intermediate CTC
- Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
- XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
- Position: Towards Responsible Evaluation for Text-to-Speech
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems
- Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
- SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
- ASTROCO: Self-Supervised Conformer-Style Transformers for Light-Curve Embeddings
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
- LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
- InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
- Retrieval Augmented Generation based context discovery for ASR
- Teffic-Audio: Tell Fact from Fiction
- Cross-attention conformer for context modeling in speech enhancement for ASR
- A perceptual similarity space for speech based on self-supervised speech representations
- FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
- CoLMbo: Speaker Language Model for Descriptive Profiling
- Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
- HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
- Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- BiLCNet : BiLSTM-Conformer Network for Encrypted Traffic Classification with 5G SA Physical Channel Records
- SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
- Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
- TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
- Interpreting the Role of Visemes in Audio-Visual Speech Recognition
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
- BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
- Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
- Pushing the Limits of End-to-End Diarization
- UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition
- FCPE: A Fast Context-based Pitch Estimation Model
- A long-form single-speaker real-time MRI speech dataset and benchmark
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
- GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
- SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- Error Analysis in a Modular Meeting Transcription System
- N-Singer: A Non-Autoregressive Korean Singing Voice Synthesis System for Pronunciation Enhancement
- Unified Learnable 2D Convolutional Feature Extraction for ASR
- Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
- Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR
- Graph Connectionist Temporal Classification for Phoneme Recognition
- XMUspeech Systems for the ASVspoof 5 Challenge
- Contextualized Token Discrimination for Speech Search Query Correction
- PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
- Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
- Human Motion Video Generation: A Survey
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
- NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
- Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
- AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
- From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
- Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
- CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
- Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
- Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
- Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
- HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- BConformeR: A Conformer Based on Mutual Sampling for Unified Prediction of Continuous and Discontinuous Antibody Binding Sites
- Pretrained Conformers for Audio Fingerprinting and Retrieval
- Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
- Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions
- A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models
- A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
- Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
- Joint decoding method for controllable contextual speech recognition based on Speech LLM
- Munsit at NADI 2025 Shared Task 2: Pushing the Boundaries of Multidialectal Arabic ASR with Weakly Supervised Pretraining and Continual Supervised Fine-tuning
- Revealing the Role of Audio Channels in ASR Performance Degradation
- DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
- Auditory Intelligence: Understanding the World Through Sound
- Score-Informed Transformer for Refining MIDI Velocity in Automatic Music Transcription
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
- SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
- Intermediate Loss Regularization for CTC-based Speech Recognition
- Learning Word-Level Confidence For Subword End-to-End ASR
- A Study on Regularization-Based Continual Learning Methods for Indic ASR
- DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
- SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
- LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
- Adaptive Knowledge Distillation for Device-Directed Speech Detection
- CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
- Nonlinear Framework for Speech Bandwidth Extension
- Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
- Self-Improvement for Audio Large Language Model using Unlabeled Speech
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Delta Keyword Transformer: Bringing Transformers to the Edge through Dynamically Pruned Multi-Head Self-Attention
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
- Unsupervised Speech Recognition
- SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
- Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
- MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio
- Data Augmentation with Locally-time Reversed Speech for Automatic Speech Recognition
- Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
- Decoupling recognition and transcription in Mandarin ASR
- Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
- Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
- Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
- Autoregressive Speech Enhancement via Acoustic Tokens
- ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
- Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
- AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
- Advancing STT for Low-Resource Real-World Speech
- Supporting SENĆOTEN Language Documentation Efforts with Automatic Speech Recognition
- DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
- The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
- AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
- Leveraging Beam Search Information for Confidence Estimation in E2E ASR
- STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
- Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
- Causal Foundation Models: Disentangling Physics from Instrument Properties
- Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
- Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
- Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study
- Prosody Labeling with Phoneme-BERT and Speech Foundation Models
- OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
- Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
- Eigenvoice Synthesis based on Model Editing for Speaker Generation
- Non-autoregressive Mandarin-English Code-switching Speech Recognition
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- Time-Masked Transformers with Lightweight Test-Time Adaptation for Neural Speech Decoding
- Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
- Automated Classification of Volcanic Earthquakes Using Transformer Encoders: Insights into Data Quality and Model Interpretability
- IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
- Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
- MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
- Rectifying Magnitude Neglect in Linear Attention
- NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
- MuteSwap: Visual-informed Silent Video Identity Conversion
- Neural Chinese silent speech recognition with facial electromyography
- A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
- Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
- WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
- Reconstructing Intelligible Speech from the Pressure Sensor Data in HVACs
- WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
- Unified Semi-Supervised Pipeline for Automatic Speech Recognition
- Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
- Accurate, fast, cheap: Choose three. Replacing Multi-Head-Attention with Bidirectional Recurrent Attention for Long-Form ASR
- Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
- Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
- Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
- Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
- Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
- Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
- Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
- Adaptive Control Attention Network for Underwater Acoustic Localization and Domain Adaptation
- Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
- TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
- Early Attentive Sparsification Accelerates Neural Speech Transcription
- Speech Recognition on TV Series with Video-guided Post-ASR Correction
- Unifying Streaming and Non-streaming Zipformer-based ASR
- End-to-end Audio-visual Speech Recognition with Conformers
- Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
- Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
- Dynamic Acoustic Model Architecture Optimization in Training for ASR
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- BUT System for the MLC-SLM Challenge
- SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
- A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations
- SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
- Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
- FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
- Regularizing Learnable Feature Extraction for Automatic Speech Recognition
- Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
- Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
- Label-Context-Dependent Internal Language Model Estimation for CTC
- Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
- Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
- Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
- Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
- Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
- LLM-based phoneme-to-grapheme for phoneme-based speech recognition
- End-to-End Diarization utilizing Attractor Deep Clustering
- Phi-Omni-ST: A multimodal language model for direct speech-to-speech translation
- A review of on-device fully neural end-to-end automatic speech recognition algorithms
- Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
- Conformer-based Ultrasound-to-Speech Conversion
- Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
- Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
- WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
- Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
- Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
- Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
- GigaAM: Efficient Self-Supervised Learner for Speech Recognition
- FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge
- Length Aware Speech Translation for Video Dubbing
- DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
- OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
- Running Conventional Automatic Speech Recognition on Memristor Hardware: A Simulated Approach
- CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
- Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
- A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
- Transformers Are Universally Consistent
- Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
- The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
- Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
- AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
- LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting
- Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
- ZIPA: A family of efficient models for multilingual phone recognition
- FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
- Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
- RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
- NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
- ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
- Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
- Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
- CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
- Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs
- PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
- Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
- Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
- Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency
- DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
- Stack Less, Repeat More: A Block Reusing Approach for Progressive Speech Enhancement
- Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
- MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
- Test-Time Adaptation with Binary Feedback
- Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
- Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain
- Large Language Models based ASR Error Correction for Child Conversations
- Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
- Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
- Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training
- Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
- In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
- ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
- Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
- The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
- QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
- Improving endpoint detection in end-to-end streaming ASR for conversational speech
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
- Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
- Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
- OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
- WIND: Accelerated RNN-T Decoding with Windowed Inference for Non-blank Detection
- The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
- Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- A Survey of Deep Learning for Complex Speech Spectrograms
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
- Unified Sparse-Matrix Representations for Diverse Neural Architectures
- Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- SpectrumFM: A Foundation Model for Intelligent Spectrum Management
- Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts
- Voice Cloning: Comprehensive Survey
- Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition
- LLMs and Speech: Integration vs. Combination
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Polynomial Mixing for Efficient Self-supervised Speech Encoders
- Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
- AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
- Pretraining Large Brain Language Model for Active BCI: Silent Speech
- A Comparative Study on Positional Encoding for Time-frequency Domain Dual-path Transformer-based Source Separation Models
- Buy versus Build an LLM: A Decision Framework for Governments
- Spatial Speech Translation: Translating Across Space With Binaural Hearables
- Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
- Multi-Task Multi-Frame Visual Piano Transcription
- Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation
- On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models
- DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
- Selective Masking Adversarial Attack on Automatic Speech Recognition Systems
- Predicting Information Pathways Across Online Communities
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Reduce and Reconstruct: ASR for Low-Resource Phonetic Languages
- Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
- CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes
- Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning
- DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
- Local Temporal Feature Enhanced Transformer with ROI-rank Based Masking for Diagnosis of ADHD
- Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
- LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models
- Summarizing Speech: A Comprehensive Survey
- RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
- Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Related