SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
2023/05/18 by Dong Zhang, Zhang, Dong, Shimin Li +11 · 123 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2305.11000
Abstract
Multi-modal large language models are regarded as a crucial step towards Artificial General Intelligence (AGI) and have garnered significant interest with the emergence of ChatGPT. However, current speech-language models typically adopt the cascade paradigm, preventing inter-modal knowledge transfer. In this paper, we propose SpeechGPT, a large language model with intrinsic cross-modal conversational abilities, capable of perceiving and generating multi-model content. With discrete speech representations, we first construct SpeechInstruct, a large-scale cross-modal speech instruction dataset. Additionally, we employ a three-stage training strategy that includes modality-adaptation pre-training, cross-modal instruction fine-tuning, and chain-of-modality instruction fine-tuning. The experimental results demonstrate that SpeechGPT has an impressive capacity to follow multi-modal human instructions and highlight the potential of handling multiple modalities with one model. Demos are shown in https://0nutation.github.io/SpeechGPT.github.io/.
Cited by
- Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
- SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Spoken Conversational Agents with Large Language Models
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- Table as a Modality for Large Language Models
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
- MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
- TASU: Text-Only Alignment for Speech Understanding
- TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
- Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
- VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- Adapting Speech Foundation Models with Large Language Models for Unified Speech Recognition
- Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
- Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- Data-Centric Lessons To Improve Speech-Language Pretraining
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
- Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
- Universal Discrete-Domain Speech Enhancement
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- Drax: Speech Recognition with Discrete Flow Matching
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
- Backdoor Attacks Against Speech Language Models
- MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
- Benchmarking Gaslighting Attacks Against Speech Large Language Models
- Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
- Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- STAR: Speech-to-Audio Generation via Representation Learning
- FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
- VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
- Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
- Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech
- Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization
- EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
- VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
- FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
- Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Group Relative Policy Optimization for Speech Recognition
- Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
- SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
- SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement
- CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
- Benchmarking Prosody Encoding in Discrete Speech Tokens
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
- MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
- UniTalker: Conversational Speech-Visual Synthesis
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
- Self-Improvement for Audio Large Language Model using Unlabeled Speech
- ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
- FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
- MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
- DIFFA: Large Language Diffusion Models Can Listen and Understand
- TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
- GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
- FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
- Step-Audio 2 Technical Report
- Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
- Autoregressive Speech Enhancement via Acoustic Tokens
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Multimodal Representation Alignment for Cross-modal Information Retrieval
- Teaching Physical Awareness to LLMs through Sounds
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
- ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition
- StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
- OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
- Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
- SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
- AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
- WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
- Aligning Spoken Dialogue Models from User Interactions
- End-to-End Spoken Grammatical Error Correction
- PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
- Watermarking Autoregressive Image Generation
- PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
- Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
- S2ST-Omni: An Efficient Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Progressive Fine-tuning
Related