Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
2025/07/24 by Samuele Cornell, Cornell, Samuele, Christoph Boeddeker +21 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2507.18161
Abstract
The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. This paper outlines the challenges' design, evaluation metrics, datasets, and baseline systems while analyzing key trends from participant submissions. From this analysis it emerges that: 1) Most participants use end-to-end (e2e) ASR systems, whereas hybrid systems were prevalent in previous CHiME challenges. This transition is mainly due to the availability of robust large-scale pre-trained models, which lowers the data burden for e2e-ASR. 2) Despite recent advances in neural speech separation and enhancement (SSE), all teams still heavily rely on guided source separation, suggesting that current neural SSE techniques are still unable to reliably deal with complex scenarios and different recording setups. 3) All best systems employ diarization refinement via target-speaker diarization techniques. Accurate speaker counting in the first diarization pass is thus crucial to avoid compounding errors and CHiME-8 DASR participants especially focused on this part. 4) Downstream evaluation via meeting summarization can correlate weakly with transcription quality due to the remarkable effectiveness of large-language models in handling errors. On the NOTSOFAR-1 scenario, even systems with over 50% time-constrained minimum permutation WER can perform roughly on par with the most effective ones (around 11%). 5) Despite recent progress, accurately transcribing spontaneous speech in challenging acoustic environments remains difficult, even when using computationally intensive system ensembles.
Citations
- UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
- Target Speaker ASR with Whisper
- Leveraging Self-Supervised Learning for Speaker Diarization
- Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
- Neural Blind Source Separation and Diarization for Distant Speech Recognition
- A Large-Scale Evaluation of Speech Foundation Models
- Rotary Position Embedding for Vision Transformer
- NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
- Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
- Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
- Profile-Error-Tolerant Target-Speaker Voice Activity Detection
- MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
- GPU-accelerated Guided Source Separation for Meeting Transcription
- Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation
- Towards a Unified Multi-Dimensional Evaluator for Text Generation
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
- End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation
- Generative Spoken Dialogue Language Modeling
- Multi-scale Speaker Diarization with Dynamic Scale Weighting
- Listen, Adapt, Better WER: Source-free Single-utterance Test-time Adaptation for Automatic Speech Recognition
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context
- Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR
- Leveraging Pretrained Models for Automatic Summarization of Doctor-Patient Conversations
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech
- SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
- SUPERB: Speech processing Universal PERformance Benchmark
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
- End-to-end speaker segmentation for overlap-aware resegmentation
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
- Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
- DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs
- WER we are and WER we think we are
- FSD50K: An Open Dataset of Human-Labeled Sound Events
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- SpecAugment on Large Scale Datasets
- Advances in Online Audio-Visual Meeting Transcription
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention\n With Dilated 1D Convolutions
- End-to-End Neural Speaker Diarization with Self-attention
- End-to-End Neural Speaker Diarization with Permutation-Free Objectives
- Optuna: A Next-generation Hyperparameter Optimization Framework
- The Second DIHARD Diarization Challenge: Dataset, task, and baselines
- Voices Obscured in Complex Environmental Settings (VOICES) corpus
- ESPnet: End-to-End Speech Processing Toolkit
- The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines
- The third ‘CHiME’ speech separation and recognition challenge: Analysis and outcomes
- Attention Is All You Need
- Sequence Transduction with Recurrent Neural Networks
- A tutorial on hidden Markov models and selected applications in speech recognition
- Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
Cited by
Related