OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
2025/05/31 by Yifan Peng, Shakeel Muhammad, Peng, Yifan +11 · 15 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2506.00338
openalex publication_date 2025/05/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a large-scale web-crawled dataset with a Creative Commons license. However, incorporating YODAS is nontrivial due to its wild nature, which introduces challenges such as incorrect language labels and audio-text misalignments. To address this, we develop a scalable data-cleaning pipeline using public toolkits, yielding a dataset with 166,000 hours of speech across 75 languages. Our new series of OWSM v4 models, trained on this curated dataset alongside existing OWSM data, significantly outperform previous versions on multilingual benchmarks. Our models even match or surpass frontier industrial models like Whisper and MMS in multiple scenarios. We will publicly release the cleaned YODAS data, pre-trained models, and all associated scripts via the ESPnet toolkit.
Citations
- MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
- MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
- On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
- OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
- OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
- YODAS: Youtube-Oriented Dataset for Audio and Speech
- Zipformer: A faster and better encoder for automatic speech recognition
- Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
- Scaling Speech Technology to 1,000+ Languages
- A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks
- Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
- Robust Speech Recognition via Large-Scale Weak Supervision
- E-Branchformer: Branchformer with Enhanced merging for speech recognition
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
- Squeezeformer: An Efficient Transformer for Automatic Speech Recognition
- FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
- Unsupervised Data Selection via Discrete Speech Representation for ASR
- The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
- SpeechBrain: A General-Purpose Speech Toolkit
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Common Voice: A Massively-Multilingual Speech Corpus
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- ESPnet: End-to-End Speech Processing Toolkit
- FastText.zip: Compressing text classification models
- Bag of Tricks for Efficient Text Classification
Cited by
Related