Common Voice: A Massively-Multilingual Speech Corpus
2019/12/13 by Rosana Ardila, Ardila, Rosana, Megan Branson +18 · 1 voice · 416 citations
Computer Science · #Artificial intelligence #Computer science #Crowdsourcing #German #Linguistics #Music and Audio Processing #Natural language processing #Speech Recognition and Synthesis #Speech analytics #Speech and Audio Processing #Speech corpus #Speech recognition #Speech synthesis #Turkish #Welsh #Word error rate #World Wide Web #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1912.06670
published in arXiv (Cornell University) (Cornell University) · Accepted to LREC 2020
openalex publication_date 2019/12/13 · arxiv created 2020/03/05 · arxiv updated 2020/03/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla's DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 +/- 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Cited by
- Efficient Multilingual ASR Finetuning via LoRA Language Experts
- MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
- Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
- FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis
- Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
- GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
- Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks
- UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
- Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
- A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
- On the Interpretability of Whisper Encodings Using Sparse Autoencoders
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
- Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
- Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
- Tracking the emergence of linguistic structure in self-supervised models learning from speech
- Greater accessibility can amplify discrimination in generative AI
- WAXAL: A Large-Scale Multilingual African Language Speech Corpus
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- Towards automating the Frenchay dysarthria assessment: Can neural phoneme posteriorgrams inform the analysis of dysarthric speech?
- Responsible Benchmarking of Fairness for Automatic Speech Recognition
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
- Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
- MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
- Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
- Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
- /UnmuteAll: Modeling Verbal Communication Patterns of Collaborative Contexts in MOBA Games
- Aliasing-Free Neural Audio Synthesis
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
- Phoneme-based speech recognition driven by large language models and sampling marginalization
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
- Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
- Segmental Attention Decoding With Long Form Acoustic Encodings
- Towards Interactive Intelligence for Digital Humans
- F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input Adaptation
- Leveraging Language Information for Target Language Extraction
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
- The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
- Qwen3.5-Omni Technical Report
- Qwen3-ASR Technical Report
- Large Speech Model Enabled Semantic Communication
- Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- Swivuriso: The South African Next Voices Multilingual Speech Dataset
- ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
- PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
- Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
- HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
- Towards Audio Token Compression in Large Audio Language Models
- Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
- Dealing with the Hard Facts of Low-Resource African NLP
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
- TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
- On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
- How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer
- Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
- IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
- CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
- Open Source State-Of-the-Art Solution for Romanian Speech Recognition
- TASU: Text-Only Alignment for Speech Understanding
- Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
- LongCat-Flash-Omni Technical Report
- Active Learning with Task-Driven Representations for Messy Pools
- Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
- Hallucination Benchmark for Speech Foundation Models
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- Interpreting the Dimensions of Speaker Embedding Space
- MPSA-DenseNet: A novel deep learning model for English accent classification
- VoiceMorph: How AI Voice Morphing Reveals the Boundaries of Auditory Self-Recognition
- A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
- The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
- FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
- M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- MLMA: Towards Multilingual ASR With Mamba-based Architectures
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- Quechua Speech Datasets in Common Voice: The Case of Puno Quechua
- video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory
- ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- LASER: An LLM-based ASR Scoring and Evaluation Rubric
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP
- Linguistically Informed Tokenization Improves ASR for Underresourced Languages
- How I Built ASR for Endangered Languages with a Spoken Dictionary
- Drax: Speech Recognition with Discrete Flow Matching
- Synthetic Audio Forensics Evaluation (SAFE) Challenge
- Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
- Reference-free automatic speech severity evaluation using acoustic unit language modelling
- XPPG-PCA: Reference-free automatic speech severity evaluation with principal components
- EuroSpeech: A Multilingual Speech Corpus
- Subjective quality evaluation of personalized own voice reconstruction systems
- Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
- Combining Knowledge Graphs and NLP to Analyze Instant Messaging Data in Criminal Investigations
- The silence of the weights: an investigation of structural pruning strategies for attention-based audio signal architectures
- The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
- Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels
- Consonant lengthening marks the beginning of words across a diverse sample of languages
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
- A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems
- Automatic Speech Recognition for Greek Medical Dictation
- AraS2P: Arabic Speech-to-Phonemes System
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
- DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
- UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
- Hi-Fi Multi-Speaker English TTS Dataset
- Weakly Supervised Phonological Features for Pathological Speech Analysis
- SwissGPC v1.0 -- The Swiss German Podcasts Corpus
- SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
- Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
- PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
- MAGIC: Multi-task Gaussian process for joint imputation and classification in healthcare time series
- Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now' in Persian
- Teffic-Audio: Tell Fact from Fiction
- Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications
- SloPalSpeech: A 2,8000-Hour Slovak Speech Corpus from Parliamentary Data
- Training Flow Matching Models with Reliable Labels via Self-Purification
- Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
- Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
- Group Relative Policy Optimization for Text-to-Speech with Large Language Models
- LOTUSDIS: A Thai far-field meeting corpus for robust conversational ASR
- DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment
- ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
- Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
- Attention-based Mixture of Experts for Robust Speech Deepfake Detection
- DIVERS-Bench: Evaluating Language Identification Across Domain Shifts and Code-Switching
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
- SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
- CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures
- VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
- Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
- HARNESS: Lightweight Distilled Arabic Speech Foundation Models
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
- TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
- QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus
- SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
- SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
- Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
- UniCoM: A Universal Code-Switching Speech Generator
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
- TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Graph Connectionist Temporal Classification for Phoneme Recognition
- DarkStream: real-time speech anonymization with low latency
- LibriQuote: A Speech Dataset of Fictional Character Utterances for Expressive Zero-Shot Speech Synthesis
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
- Group Relative Policy Optimization for Speech Recognition
- Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
- Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
- Analysing the Language of Neural Audio Codecs
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
- Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
- Towards Improved Speech Recognition through Optimized Synthetic Data Generation
- Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
- Beyond Transcription: Mechanistic Interpretability in ASR
- TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
- Efficient neural encoding as revealed by bilingualism
- Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
- Cryfish: On deep audio analysis with Large Language Models
- Lossless data compression by large models
- Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
- Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
- Out of the Box, into the Clinic? Evaluating State-of-the-Art ASR for Clinical Applications for Older Adults
- DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
- Scalable Controllable Accented TTS
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
- Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
- NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
- Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
- Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy
- Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
- SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
- Advancing the Foundation Model for Music Understanding
- Evaluating and Improving the Robustness of Speech Command Recognition Models to Noise and Distribution Shifts
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
- Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
- Self-Improvement for Audio Large Language Model using Unlabeled Speech
- Voxtral
- Improving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Data
- Comparison of Knowledge Distillation Methods for Low-complexity Multi-microphone Speech Enhancement using the FT-JNF Architecture
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- Synthetic Data Generation for Phrase Break Prediction with Large Language Model
- The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
- WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes
- DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
- FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
- MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio
- Self-Supervised Representations Improve End-to-End Speech Translation
- Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
- Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
- Synthetic Voice Data for Automatic Speech Recognition in African Languages
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Spoken language identification: An overview of past and present research trends
- Exploiting Pre-Trained ASR Models for Alzheimer's Disease Recognition Through Spontaneous Speech
- P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
- WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
- A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
- The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
- Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
- Unlocking Speech Instruction Data Potential with Query Rewriting
- ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition
- Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
- Leveraging Beam Search Information for Confidence Estimation in E2E ASR
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
- Differentiable Reward Optimization for LLM based TTS system
- Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
- Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
- Word stress in self-supervised speech models: A cross-linguistic comparison
- Self-supervised learning of speech representations with Dutch archival data
- MMMOS: Multi-domain Multi-axis Audio Quality Assessment
- Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
- A Cookbook for Community-driven Data Collection of Impaired Speech in LowResource Languages
- Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
- NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
- Dataset to publication: "Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?"
- WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
- Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
- SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition
- Unified Semi-Supervised Pipeline for Automatic Speech Recognition
- Hybrid Deep Learning and Signal Processing for Arabic Dialect Recognition in Low-Resource Settings
- Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
- Context Biasing for Pronunciations-Orthography Mismatch in Automatic Speech Recognition
- Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
- USAD: Universal Speech and Audio Representation via Distillation
- Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings
- Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
- Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
- Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
- Howl: A Deployed, Open-Source Wake Word Detection System
- Weight Factorization and Centralization for Continual Learning in Speech Recognition
- End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
- DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction
- Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
- CMU's IWSLT 2025 Simultaneous Speech Translation System
- E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
- Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
- Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
- Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
- Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
- Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
- WAKE: Watermarking Audio with Key Enrichment
- Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition
- LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
- EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
- LLM-based phoneme-to-grapheme for phoneme-based speech recognition
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- The mutual exclusivity bias of bilingual visually grounded speech models
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
- Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
- A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
- Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
- On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
- SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
- Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
- Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
- Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
- Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data
- WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
- Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
- Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
- Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
- GigaAM: Efficient Self-Supervised Learner for Speech Recognition
- Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition
- LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
- XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
- Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
- OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
- Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
- MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
- Improving Language and Modality Transfer in Translation by Character-level Modeling
- ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
- Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
- FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System
- SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
- Children's Voice Privacy: First Steps And Emerging Challenges
- The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
- Interspeech 2025 URGENT Speech Enhancement Challenge
- ZIPA: A family of efficient models for multilingual phone recognition
- StressTest: Can YOUR Speech LM Handle the Stress?
- FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
- Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
- Voice Adaptation for Swiss German
- Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
- ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
- Simulating Early Phonetic and Word Learning Without Linguistic Categories
- Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
- GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
- PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead
- The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages
- DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
- CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
- Topic Model Robustness to Automatic Speech Recognition Errors in Podcast Transcripts
- LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
- VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
- TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
- Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
- Speechless: Speech Instruction Training Without Speech for Low Resource Languages
- LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
- HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification
- On the reliability of feature attribution methods for speech classification
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
- VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
- Word Level Timestamp Generation for Automatic Speech Recognition and Translation
- Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
- A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
- Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
- Fast, Not Fancy: Rethinking G2P with Rich Data and Rule-Based Models
- LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
- Joint gender and age estimation based on speech signals using x-vectors and transfer learning
- BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake Audio Dataset
- UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
- MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
- Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
- Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
- Automatic Speech Recognition Benchmark for Air-Traffic Communications
- Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients
- Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations
- StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
- Uncertainty Estimation in Autoregressive Structured Prediction
- Swiss Parliaments Corpus, an Automatically Aligned Swiss German Speech to Standard German Text Corpus
- Voice Cloning: Comprehensive Survey
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- A GAN-based Approach for Mitigating Inference Attacks in Smart Home Environment
- Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
- [b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
- BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
- IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
- Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
- "Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
- SpeechMapper: Speech-to-text Embedding Projector for LLMs
- Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
- BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
- Kimi-Audio Technical Report
- MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
- Affect Models Have Weak Generalizability to Atypical Speech
- Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition
- DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
- Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning
- Poem Meter Classification of Recited Arabic Poetry: Integrating High-Resource Systems for a Low-Resource Task
- Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
- Position: The Most Expensive Part of an LLM should be its Training Data
- How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs
- Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
- Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
- Spatial Audio Processing with Large Language Model on Wearable Devices
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- List of datasets for machine-learning research [wikipedia]
Discussions
Related