Multimodal Chain-of-Thought Reasoning in Language Models
2023/02/02 by Zhuosheng Zhang, Aston Zhang, Zhang, Zhuosheng +9 · 1 voice · 132 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.2302.00923
openalex publication_date 2023/02/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) have shown impressive performance on complex reasoning by leveraging chain-of-thought (CoT) prompting to generate intermediate reasoning chains as the rationale to infer the answer. However, existing CoT studies have primarily focused on the language modality. We propose Multimodal-CoT that incorporates language (text) and vision (images) modalities into a two-stage framework that separates rationale generation and answer inference. In this way, answer inference can leverage better generated rationales that are based on multimodal information. Experimental results on ScienceQA and A-OKVQA benchmark datasets show the effectiveness of our proposed approach. With Multimodal-CoT, our model under 1 billion parameters achieves state-of-the-art performance on the ScienceQA benchmark. Our analysis indicates that Multimodal-CoT offers the advantages of mitigating hallucination and enhancing convergence speed. Code is publicly available at https://github.com/amazon-science/mm-cot.
Cited by
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
- ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
- RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
- LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- Visual Access Boundaries in Vision-Language Model Reasoning
- Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
- ICLR: In-Context Imitation Learning with Visual Reasoning
- Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
- Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- SuperCLIP: CLIP with Simple Classification Supervision
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Rethinking Chain-of-Thought Reasoning for Videos
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- OneThinker: All-in-one Reasoning Model for Image and Video
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- VACoT: Rethinking Visual Data Augmentation with VLMs
- SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Personalized Reward Modeling for Text-to-Image Generation
- Step-Audio-R1 Technical Report
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
- Remodeling Semantic Relationships in Vision-Language Fine-Tuning
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
- Latent Chain-of-Thought for Visual Reasoning
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- 3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- A Survey on Agentic Multimodal Large Language Models
- Evaluating Language Models' Evaluations of Games
- Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
- ContextNav: Towards Agentic Multimodal In-Context Learning
- MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- Latent Visual Reasoning
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Planning with Unified Multimodal Models
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections
- UniAPO: Unified Multimodal Automated Prompt Optimization
- Instant Preference Alignment for Text-to-Image Diffusion Models
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- LIMI: Less is More for Agency
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
- Chain-of-Thought Re-ranking for Image Retrieval Tasks
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- A QoE-Driven Personalized Incentive Mechanism Design for AIGC Services in Resource-Constrained Edge Networks
- FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
- Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Controlling Multimodal LLMs via Reward-guided Decoding
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- MolmoAct: Action Reasoning Models that can Reason in Space
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Can Large Vision-Language Models Understand Multimodal Sarcasm?
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
- Language as Cost: Proactive Hazard Mapping using VLM for Robot Navigation
- ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
Discussions
Related