Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
2025/03/16 by Yaoting Wang, Wang, Yaoting, Shengqiong Wu +11 · 1 voice · 52 citations
#cs.CV
paper · pdf · doi:10.48550/arxiv.2503.12605
Abstract
By extending the advantage of chain-of-thought (CoT) reasoning in human-like step-by-step processes to multimodal contexts, multimodal CoT (MCoT) reasoning has recently garnered significant research attention, especially in the integration with multimodal large language models (MLLMs). Existing MCoT studies design various methodologies and innovative reasoning paradigms to address the unique challenges of image, video, speech, audio, 3D, and structured data across different modalities, achieving extensive success in applications such as robotics, healthcare, autonomous driving, and multimodal generation. However, MCoT still presents distinct challenges and opportunities that require further focus to ensure consistent thriving in this field, where, unfortunately, an up-to-date review of this domain is lacking. To bridge this gap, we present the first systematic survey of MCoT reasoning, elucidating the relevant foundational concepts and definitions. We offer a comprehensive taxonomy and an in-depth analysis of current methodologies from diverse perspectives across various application scenarios. Furthermore, we provide insights into existing challenges and future research directions, aiming to foster innovation toward multimodal AGI.
Cited by
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
- CogDoc: Towards Unified thinking in Documents
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
- LaRe: Latent Refocusing for Multimodal Reasoning
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement Learning
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- A Survey on Agentic Multimodal Large Language Models
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models
- MuSLR: Multimodal Symbolic Logical Reasoning
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
- Fast Thinking for Large Language Models
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
- Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
- Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
- RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- MolmoAct: Action Reasoning Models that can Reason in Space
- Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models
- CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- A2R2: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
Discussions
Related