Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
2025/10/26 by Canxiang Yan, Yan, Canxiang, Dawei Huang +41 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Music Technology and Sound Studies
paper · pdf · doi:10.48550/arxiv.2511.05516
Abstract
Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form editing. To solve this challenge, we introduce a novel framework that unifies speech understanding, generation, and editing. The core of our unified model is a unified continuous speech tokenizer MingTok-Audio, the first continuous tokenizer to effectively integrate semantic and acoustic features, which makes it suitable for both understanding and generation tasks. Based on this unified continuous audio tokenizer, we developed the speech language model Ming-UniAudio, which achieved a balance between generation and understanding capabilities. Ming-UniAudio sets new state-of-the-art (SOTA) records on 8 out of 12 metrics on the ContextASR benchmark. Notably, for Chinese voice cloning, it achieves a highly competitive Seed-TTS-WER of 0.95. Leveraging this foundational model, we further trained a dedicated speech editing model Ming-UniAudio-Edit, the first speech language model that enables universal, free-form speech editing guided solely by natural language instructions, handling both semantic and acoustic modifications without timestamp condition. To rigorously assess the editing capability and establish a foundation for future research, we introduce Ming-Freeform-Audio-Edit, the first comprehensive benchmark tailored for instruction-based free-form speech editing, featuring diverse scenarios and evaluation dimensions spanning semantic correctness, acoustic quality, and instruction alignment. We open-sourced the continuous audio tokenizer, the unified foundational model, and the free-form instruction-based editing model to facilitate the development of unified audio understanding, generation, and manipulation.
Citations
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Kimi-Audio Technical Report
- Qwen2.5-Omni Technical Report
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- TS3-Codec: Transformer-Based Simple Streaming Single Codec
- GPT-4o System Card
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
- Moshi: a speech-text foundation model for real-time dialogue
- BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
- FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
- Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
- SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
- SlideSpeech: A Large-Scale Slide-Enriched Audio-Visual Corpus
- Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
- Inter-SubNet: Speech Enhancement with Subband Interaction
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Robust Speech Recognition via Large-Scale Weak Supervision
- High Fidelity Neural Audio Compression
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
- Conditional Diffusion Probabilistic Model for Speech Enhancement
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- EdiTTS: Score-based Editing for Controllable Text-to-Speech
- DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- SoundStream: An End-to-End Neural Audio Codec
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
- Librispeech Transducer Model with Internal Language Model Prior Correction
- SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
- AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
- CoVoST 2 and Massively Multilingual Speech-to-Text Translation
- Language Models are Few-Shot Learners
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
- CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- MUSAN: A Music, Speech, and Noise Corpus
Cited by
Related