ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
2025/07/08 by He Wang, Linhan Ma, Wang, He +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2507.05727
openalex publication_date 2025/07/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for accurately recognizing named entities across diverse domains. For instance, drug and treatment names in medicine or specialized technical terms in engineering. Recent breakthroughs in Large Language Models (LLMs) and corresponding Large Audio Language Models (LALMs) have markedly enhanced the visibility of advanced context modeling and general artificial intelligence capabilities. Leveraging LLMs, we envision a unified system capable of robust speech recognition across diverse real-world domains, yet existing benchmarks are inadequate for evaluating this objective. To address this gap, we propose ContextASR-Bench: a comprehensive, large-scale benchmark designed to assess the linguistic competence of ASR systems using corpora that feature numerous named entities across multiple domains. It encompasses up to 40,000 data entries with more than 300,000 named entities across over 10 domains. Beyond the audio and its transcription, each sample provides the domain it belongs to and a list of named entities it contains, which are referred to as the context. Based on this, we introduce three evaluation modes to assess how effectively models can exploit such context to improve ASR accuracy. Extensive evaluation on ContextASR-Bench highlights that LALMs outperform conventional ASR models by a large margin thanks to the strong world knowledge and context modeling of LLMs, yet there remains ample room for further improvement. The dataset and evaluation code have been released.
Citations
- Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
- AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
- NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
- Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
- A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
- SlideSpeech: A Large-Scale Slide-Enriched Audio-Visual Corpus
- Graph Neural Networks for Contextual ASR with the Tree-Constrained Pointer Generator
- FunASR: A Fundamental End-to-End Speech Recognition Toolkit
- Robust Speech Recognition via Large-Scale Weak Supervision
- FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
- A Benchmark for Automatic Medical Consultation System: Frameworks, Tasks and Datasets
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
- DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
- SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
- DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressors
- Measuring Massive Multitask Language Understanding
- The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for Chinese
- Common Voice: A Massively-Multilingual Speech Corpus
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- Chinese NER Using Lattice LSTM
- Exploring Architectures, Data and Units For Streaming End-to-End Speech Recognition with RNN-Transducer
- A Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
- THCHS-30 : A Free Chinese Speech Corpus
Cited by
Related