When to Think and When to Look: Uncertainty-Guided Lookback
2025/11/19 by Bi, Jing, Bellos, Filippos, Guo, Junjia +8 · 2 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.15613
Abstract
Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how thinking actually affects visual reasoning. We provide the first such analysis with a large scale, controlled comparison of thinking for LVLMs, evaluating ten variants from the InternVL3.5 and Qwen3-VL families on MMMU-val under generous token budgets and multi pass decoding. We show that more thinking is not always better; long chains often yield long wrong trajectories that ignore the image and underperform the same models run in standard instruct mode. A deeper analysis reveals that certain short lookback phrases, which explicitly refer back to the image, are strongly enriched in successful trajectories and correlate with better visual grounding. Building on this insight, we propose uncertainty guided lookback, a training free decoding strategy that combines an uncertainty signal with adaptive lookback prompts and breadth search. Our method improves overall MMMU performance, delivers the largest gains in categories where standard thinking is weak, and outperforms several strong decoding baselines, setting a new state of the art under fixed model families and token budgets. We further show that this decoding strategy generalizes, yielding consistent improvements on five additional benchmarks, including two broad multimodal suites and math focused visual reasoning datasets.
Citations
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
- Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Deep Think with Confidence
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
- Reasoning Models Don't Always Say What They Think
- A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
- Dynamic Early Exit in Reasoning Models
- Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
- Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
- Grounded Chain-of-Thought for Multimodal Large Language Models
- VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
- Qwen2.5-VL Technical Report
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- Demystifying Long Chain-of-Thought Reasoning in LLMs
- Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MMBench: Is Your Multi-modal Model an All-around Player?
- Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning
Cited by
Related