M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
2024/05/26 by Qiguang Chen, Chen, Qiguang, Libo Qin +9 · 65 citations
Computer Science · Physics and Astronomy · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Complex Network Analysis Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2405.16473
openalex publication_date 2024/05/26 · openalex created_date 2024/05/29 · openalex updated_date 2026/07/28
Abstract
Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-by-step reasoning, which gains increasing attention. Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT. Motivated by this, we introduce a novel benchmark (M3CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT. Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs). In addition, we highlight that the current VLLMs still struggle to correctly reason in M3CoT and there remains a large gap between existing VLLMs and human performance in M3CoT, despite their superior results on previous MCoT benchmarks. To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT. We hope that M3CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research.
Cited by
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- A Survey on Agentic Multimodal Large Language Models
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- A Survey of Deep Learning for Geometry Problem Solving
- ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
- Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
- Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
- Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
- RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
- Visual Planning: Let's Think Only with Images
- Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
- Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach
- MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
Related