Multimodal Chain-of-Thought Reasoning in Language Models
2023/02/02 by Zhuosheng Zhang, Aston Zhang, Zhang, Zhuosheng +9 · 1 voice · 227 citations
Computer Science · #Artificial intelligence #Benchmark (surveying) #Computer science #Inference #Language model #Leverage (statistics) #Machine learning #Modalities #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Topic Modeling #cs.AI #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.2302.00923
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/02/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) have shown impressive performance on complex reasoning by leveraging chain-of-thought (CoT) prompting to generate intermediate reasoning chains as the rationale to infer the answer. However, existing CoT studies have primarily focused on the language modality. We propose Multimodal-CoT that incorporates language (text) and vision (images) modalities into a two-stage framework that separates rationale generation and answer inference. In this way, answer inference can leverage better generated rationales that are based on multimodal information. Experimental results on ScienceQA and A-OKVQA benchmark datasets show the effectiveness of our proposed approach. With Multimodal-CoT, our model under 1 billion parameters achieves state-of-the-art performance on the ScienceQA benchmark. Our analysis indicates that Multimodal-CoT offers the advantages of mitigating hallucination and enhancing convergence speed. Code is publicly available at https://github.com/amazon-science/mm-cot.
Cited by
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
- ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
- RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
- LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- Visual Access Boundaries in Vision-Language Model Reasoning
- Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
- ICLR: In-Context Imitation Learning with Visual Reasoning
- Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
- Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- SuperCLIP: CLIP with Simple Classification Supervision
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Rethinking Chain-of-Thought Reasoning for Videos
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- OneThinker: All-in-one Reasoning Model for Image and Video
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- VACoT: Rethinking Visual Data Augmentation with VLMs
- SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Personalized Reward Modeling for Text-to-Image Generation
- Step-Audio-R1 Technical Report
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
- Remodeling Semantic Relationships in Vision-Language Fine-Tuning
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
- Latent Chain-of-Thought for Visual Reasoning
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- 3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- A Survey on Agentic Multimodal Large Language Models
- Evaluating Language Models' Evaluations of Games
- Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
- ContextNav: Towards Agentic Multimodal In-Context Learning
- MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- Latent Visual Reasoning
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Planning with Unified Multimodal Models
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections
- UniAPO: Unified Multimodal Automated Prompt Optimization
- Instant Preference Alignment for Text-to-Image Diffusion Models
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- LIMI: Less is More for Agency
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
- Chain-of-Thought Re-ranking for Image Retrieval Tasks
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- A QoE-Driven Personalized Incentive Mechanism Design for AIGC Services in Resource-Constrained Edge Networks
- FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
- Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Controlling Multimodal LLMs via Reward-guided Decoding
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- MolmoAct: Action Reasoning Models that can Reason in Space
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
- Can Large Vision-Language Models Understand Multimodal Sarcasm?
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
- Language as Cost: Proactive Hazard Mapping using VLM for Robot Navigation
- ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
- Position: Reasoning After Perception Means Reasoning Without Vision
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- Survey of GenAI for Automotive Software Development: From Requirements to Executable Code
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- Introspection of Thought Helps AI Agents
- Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
- Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata
- MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- Spatial transcriptomics AI agent charts hPSC-pancreas maturation in vivo
- Empowering Small VLMs to Think with Dynamic Memorization and Exploration
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
- MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought
- PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
- DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object Segmentation
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
- Generating Automotive Code: Large Language Models for Software Development and Verification in Safety-Critical Systems
- Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
- Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
- Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- SiLVR: A Simple Language-based Video Reasoning Framework
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- Are MLMs Trapped in the Visual Room?
- DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
- Thinking with Generated Images
- Reinforced Reasoning for Embodied Planning
- Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection
- PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- Douyin Multimodal Embedding Model Technical Report
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
- Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
- VLM-R3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
- Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
- CAMA: Enhancing Multimodal In-Context Learning with Context-Aware Modulated Attention
- GRIT: Teaching MLLMs to Think with Images
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
- Visual Planning: Let's Think Only with Images
- StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
- CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing
- Advancing AI Research Assistants with Expert-Involved Learning
- ReLI: A Language-Agnostic Approach to Human-Robot Interaction
- SSRLBot: Designing and Developing a Large Language Model-based Agent using Socially Shared Regulated Learning
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
- Lightweight Visual Reasoning for Socially-Aware Robots
- Generative Visual Code Mobile World Models
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
- Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
- Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling
Discussions
Related