Common Voice: A Massively-Multilingual Speech Corpus
2019/12/13 by Rosana Ardila, Ardila, Rosana, Megan Branson +18 · 1 voice · 224 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1912.06670
openalex publication_date 2019/12/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla's DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 +/- 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Cited by
- MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
- Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
- FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis
- Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
- GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
- Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks
- UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
- Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
- A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
- On the Interpretability of Whisper Encodings Using Sparse Autoencoders
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
- Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
- Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
- Tracking the emergence of linguistic structure in self-supervised models learning from speech
- Greater accessibility can amplify discrimination in generative AI
- WAXAL: A Large-Scale Multilingual African Language Speech Corpus
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- Towards automating the Frenchay dysarthria assessment: Can neural phoneme posteriorgrams inform the analysis of dysarthric speech?
- Responsible Benchmarking of Fairness for Automatic Speech Recognition
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
- Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
- MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
- Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
- Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
- /UnmuteAll: Modeling Verbal Communication Patterns of Collaborative Contexts in MOBA Games
- Aliasing-Free Neural Audio Synthesis
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
- Phoneme-based speech recognition driven by large language models and sampling marginalization
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
- Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
- Segmental Attention Decoding With Long Form Acoustic Encodings
- Towards Interactive Intelligence for Digital Humans
- F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input Adaptation
- Leveraging Language Information for Target Language Extraction
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
- The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
- Qwen3.5-Omni Technical Report
- Qwen3-ASR Technical Report
- Large Speech Model Enabled Semantic Communication
- Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- Swivuriso: The South African Next Voices Multilingual Speech Dataset
- ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
- PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
- Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
- HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
- Towards Audio Token Compression in Large Audio Language Models
- Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
- Dealing with the Hard Facts of Low-Resource African NLP
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
- TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
- On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
- How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer
- Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
- IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
- CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
- Open Source State-Of-the-Art Solution for Romanian Speech Recognition
- TASU: Text-Only Alignment for Speech Understanding
- Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
- LongCat-Flash-Omni Technical Report
- Active Learning with Task-Driven Representations for Messy Pools
- Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
- Hallucination Benchmark for Speech Foundation Models
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- Interpreting the Dimensions of Speaker Embedding Space
- MPSA-DenseNet: A novel deep learning model for English accent classification
- VoiceMorph: How AI Voice Morphing Reveals the Boundaries of Auditory Self-Recognition
- A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
- The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
- FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
- M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- MLMA: Towards Multilingual ASR With Mamba-based Architectures
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- Quechua Speech Datasets in Common Voice: The Case of Puno Quechua
- video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory
- ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- LASER: An LLM-based ASR Scoring and Evaluation Rubric
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP
- Linguistically Informed Tokenization Improves ASR for Underresourced Languages
- How I Built ASR for Endangered Languages with a Spoken Dictionary
- Drax: Speech Recognition with Discrete Flow Matching
- Synthetic Audio Forensics Evaluation (SAFE) Challenge
- Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
- Reference-free automatic speech severity evaluation using acoustic unit language modelling
- XPPG-PCA: Reference-free automatic speech severity evaluation with principal components
- EuroSpeech: A Multilingual Speech Corpus
- Subjective quality evaluation of personalized own voice reconstruction systems
- Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
- Combining Knowledge Graphs and NLP to Analyze Instant Messaging Data in Criminal Investigations
- The silence of the weights: an investigation of structural pruning strategies for attention-based audio signal architectures
- The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
- Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels
- Consonant lengthening marks the beginning of words across a diverse sample of languages
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
- A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems
- Automatic Speech Recognition for Greek Medical Dictation
- AraS2P: Arabic Speech-to-Phonemes System
- ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
- DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
- UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
- Hi-Fi Multi-Speaker English TTS Dataset
- Weakly Supervised Phonological Features for Pathological Speech Analysis
- SwissGPC v1.0 -- The Swiss German Podcasts Corpus
- SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
- Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
- PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
- MAGIC: Multi-task Gaussian process for joint imputation and classification in healthcare time series
- Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now' in Persian
- Teffic-Audio: Tell Fact from Fiction
- Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications
- SloPalSpeech: A 2,8000-Hour Slovak Speech Corpus from Parliamentary Data
- Training Flow Matching Models with Reliable Labels via Self-Purification
- Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
- Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
- Group Relative Policy Optimization for Text-to-Speech with Large Language Models
- LOTUSDIS: A Thai far-field meeting corpus for robust conversational ASR
- DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment
- ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
- Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
- Attention-based Mixture of Experts for Robust Speech Deepfake Detection
- DIVERS-Bench: Evaluating Language Identification Across Domain Shifts and Code-Switching
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
- SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
- CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures
- VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
- Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
- HARNESS: Lightweight Distilled Arabic Speech Foundation Models
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
- TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
- QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic\n Speech Corpus
- SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
- SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
- Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
- UniCoM: A Universal Code-Switching Speech Generator
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
- TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Graph Connectionist Temporal Classification for Phoneme Recognition
- DarkStream: real-time speech anonymization with low latency
- LibriQuote: A Speech Dataset of Fictional Character Utterances for Expressive Zero-Shot Speech Synthesis
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
- Group Relative Policy Optimization for Speech Recognition
- Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
- Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
- Analysing the Language of Neural Audio Codecs
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
- Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
- Towards Improved Speech Recognition through Optimized Synthetic Data Generation
- Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
- Beyond Transcription: Mechanistic Interpretability in ASR
- TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
- Efficient neural encoding as revealed by bilingualism
- Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
- Cryfish: On deep audio analysis with Large Language Models
- Lossless data compression by large models
- Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
- Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
- Out of the Box, into the Clinic? Evaluating State-of-the-Art ASR for Clinical Applications for Older Adults
- DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
- Scalable Controllable Accented TTS
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
- NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
- Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
- Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
- SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
- Advancing the Foundation Model for Music Understanding
- Evaluating and Improving the Robustness of Speech Command Recognition Models to Noise and Distribution Shifts
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
- Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
- List of datasets for machine-learning research [wikipedia]
Discussions
Related