Efficient Scaling for LLM-based ASR
2025/08/06 by Mu, Bingshen, Shao, Yiwen, Wei, Kun +2 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2508.04096
Abstract
Large language model (LLM)-based automatic speech recognition (ASR) achieves strong performance but often incurs high computational costs. This work investigates how to obtain the best LLM-ASR performance efficiently. Through comprehensive and controlled experiments, we find that pretraining the speech encoder before integrating it with the LLM leads to significantly better scaling efficiency than the standard practice of joint post-training of LLM-ASR. Based on this insight, we propose a new multi-stage LLM-ASR training strategy, EFIN: Encoder First Integration. Among all training strategies evaluated, EFIN consistently delivers better performance (relative to 21.1% CERR) with significantly lower computation budgets (49.9% FLOPs). Furthermore, we derive a scaling law that approximates ASR error rates as a computation function, providing practical guidance for LLM-ASR scaling.
Citations
- Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
- Qwen3 Technical Report
- OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
- FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
- HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
- Advancing Multi-talker ASR Performance with Large Language Models
- The Llama 3 Herd of Models
- Qwen2-Audio Technical Report
- MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
- Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
- Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
- Qwen Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Scaling Laws for Discriminative Speech Recognition Rescoring Models
- AudioPaLM: A Large Language Model That Can Speak and Listen
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Scaling Laws for Multilingual Neural Machine Translation
- Robust Speech Recognition via Large-Scale Weak Supervision
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
- Scaling Laws for Neural Machine Translation
- LoRA: Low-Rank Adaptation of Large Language Models
- Scaling Laws for Acoustic Models
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- Deep Learning Scaling is Predictable, Empirically
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- Attention Is All You Need
- End-to-End Attention-based Large Vocabulary Speech Recognition
- Listen, Attend and Spell
- EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding
- Sequence Transduction with Recurrent Neural Networks
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
Cited by
Related