Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
2025/11/21 by Endo, Mark, Serena Yeung, Yeung-Levy, Serena
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2511.17487
openalex publication_date 2025/11/21 · openalex created_date 2025/11/25 · openalex updated_date 2026/07/28
Abstract
Scaling up multimodal models has enabled remarkable advances in visual understanding and reasoning, but practical demands call for smaller, efficient systems. In this work, we conduct a principled analysis of downscaling intelligence in multimodal models, examining how reduced large language model (LLM) capacity affects multimodal capabilities. Our initial findings reveal an interesting trend: LLM downscaling disproportionately affects visual capabilities, rather than abilities inherited from the LLM. We then examine whether this drop mainly reflects the expected decline in visual reasoning or a more fundamental loss of perceptual abilities. Isolating the effect of LLM downscaling on perception, we find performance still drops sharply, often matching or exceeding the impact on reasoning. To address this bottleneck, we introduce visual extraction tuning, which explicitly trains the model to extract instruction-relevant visual details consistently across tasks. With these extracted visual details, we then apply step-by-step reasoning to generate answers. Together, these components form our Extract+Think approach, setting a new standard for efficiency and performance in this space.
Citations
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- Qwen3 Technical Report
- Scaling Laws for Native Multimodal Models
- SmolVLM: Redefining small and efficient multimodal models
- Gemma 3 Technical Report
- On the Perception Bottleneck of VLMs for Chart Understanding
- Qwen2.5-VL Technical Report
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Reasoning Limitations of Multimodal Large Language Models. A Case Study of Bongard Problems
- Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
- GPT-4o System Card
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LLaVA-OneVision: Easy Visual Task Transfer
- The Llama 3 Herd of Models
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
- Why are Visually-Grounded Language Models Bad at Image Classification?
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- InternLM2 Technical Report
- PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
- CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
- VISION Datasets: A Benchmark for Vision-based InduStrial InspectiON
- Sigmoid Loss for Language Image Pre-Training
- The Quantization Model of Neural Scaling
- Scaling Laws for Generative Mixed-Modal Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Large Language Models are Zero-Shot Reasoners
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- WebQA: Multihop and Multimodal QA
- DocVQA: A Dataset for VQA on Document Images
- Neural Naturalist: Generating Fine-Grained Image Comparisons
- Expressing Visual Relationships via Language
- Towards VQA Models That Can Read
- RAVEN: A Dataset for Relational and Analogical Visual rEasoNing
- StoryGAN: A Sequential Conditional GAN for Story Visualization
- RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
- Learning to Describe Differences Between Pairs of Similar Images
- Imagine This! Scripts to Compositions to Videos
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Visual Storytelling
- Generation and Comprehension of Unambiguous Object Descriptions
Related