Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
2025/10/02 by Siddhant Arora, Jinchuan Tian, Arora, Siddhant +11
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multi-Agent Systems and Negotiation #Robotics and Automated Systems #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2510.02066
openalex publication_date 2025/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Abstract
Most end-to-end (E2E) spoken dialogue systems (SDS) rely on voice activity detection (VAD) for turn-taking, but VAD fails to distinguish between pauses and turn completions. Duplex SDS models address this by predicting output continuously, including silence tokens, thus removing the need for explicit VAD. However, they often have complex dual-channel architecture and lag behind cascaded models in semantic reasoning. To overcome these challenges, we propose SCoT: a Streaming Chain-of-Thought (CoT) framework for Duplex SDS, alternating between processing fixed-duration user input and generating responses in a blockwise manner. Using frame-level alignments, we create intermediate targets-aligned user transcripts and system responses for each block. Experiments show that our approach produces more coherent and interpretable responses than existing duplex methods while supporting lower-latency and overlapping interactions compared to turn-by-turn systems.
Citations
- Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
- OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Qwen2.5-Omni Technical Report
- ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
- ESPnet-SpeechLM: An Open Speech Language Model Toolkit
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
- WavChat: A Survey of Spoken Dialogue Models
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
- ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
- Enabling Real-Time Conversations with Minimal Training Costs
- Moshi: a speech-text foundation model for real-time dialogue
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Language Model Can Listen While Speaking
- Towards Robust Speech Representation Learning for Thousands of Languages
- Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
- A Full-duplex Speech Dialogue Scheme Based On Large Language Models
- Spirit LM: Interleaved Spoken and Written Language Model
- SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
- Simple and Controllable Music Generation
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- GPT-4 Technical Report
- Robust Speech Recognition via Large-Scale Weak Supervision
- Automatic Chain of Thought Prompting in Large Language Models
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- Generative Spoken Dialogue Language Modeling
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- ESPnet-SLU: Advancing Spoken Language Understanding through ESPnet
- Language Models are Few-Shot Learners
- ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
- Investigating Speech Features for Continuous Turn-Taking Prediction Using LSTMs
- ESPnet: End-to-End Speech Processing Toolkit
Cited by
Related