BLINK: Multimodal Large Language Models Can See but Not Perceive
2024/04/18 by Xingyu Fu, Fu, Xingyu, Yushi Hu +18 · 1 voice · 213 citations
Computer Science · Psychology · #Cognitive psychology #Computer science #Linguistics #Multimodal therapy #Natural Language Processing Techniques #Philosophy #Psychology #Psychotherapist #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2404.12390
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/04/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning). However, we find these perception-demanding tasks cast significant challenges for current multimodal LLMs because they resist mediation through natural language. Blink reformats 14 classic computer vision tasks into 3,807 multiple-choice questions, paired with single or multiple images and visual prompting. While humans get 95.70% accuracy on average, Blink is surprisingly challenging for existing multimodal LLMs: even the best-performing GPT-4V and Gemini achieve accuracies of 51.26% and 45.72%, only 13.17% and 7.63% higher than random guessing, indicating that such perception abilities have not "emerged" yet in recent multimodal LLMs. Our analysis also highlights that specialist CV models could solve these problems much better, suggesting potential pathways for future improvements. We believe Blink will stimulate the community to help multimodal LLMs catch up with human-level visual perception.
Cited by
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
- SceneActBench: Can Agents Act on the 3D Scenes They See?
- SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning
- Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
- STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
- RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
- LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
- SportD: Can VLMs Physically Strategize?
- An Exam for Active Observers
- TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
- DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
- Cross-Modal Taxonomic Generalization in (Vision-) Language Models
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- LanteRn: Latent Visual Structured Reasoning
- Latent Implicit Visual Reasoning
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- How Much 3D Do Video Foundation Models Encode?
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- Visually Prompted Benchmarks Are Surprisingly Fragile
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Towards Lossless Ultimate Vision Token Compression for VLMs
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- Towards Cross-View Point Correspondence in Vision-Language Models
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Jina-VLM: Small Multilingual Vision Language Model
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
- Qwen3-VL Technical Report
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Vision-Language Memory for Spatial Reasoning
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- Understanding Task Transfer in Vision-Language Models
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
- SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- iFlyBot-VLM Technical Report
- Visual Spatial Tuning
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- NVIDIA Nemotron Nano V2 VL
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- LongCat-Flash-Omni Technical Report
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Mitigating Modal Imbalance in Multimodal Reasoning
- Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning
- Multimodal Language Models Cannot Spot Spatial Inconsistencies
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- [De|Re]constructing VLMs' Reasoning in Counting
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- Visual Representations inside the Language Model
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Apriel-1.5-15b-Thinker
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- Latent Visual Reasoning
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- Human-like Navigation in a World Built for Humans
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- 3D Aware Region Prompted Vision Language Model
- Reconstruction Alignment Improves Unified Multimodal Models
- Reinforced Visual Perception with Tools
- Kwai Keye-VL 1.5 Technical Report
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
- Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- Thyme: Think Beyond Images
- Ovis2.5 Technical Report
- A Study of Commonsense Reasoning over Visual Object Properties
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- Position: Reasoning After Perception Means Reasoning Without Vision
- ERNIE 5.0 Technical Report
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- Kwai Keye-VL Technical Report
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- Synthetic Visual Genome
- Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- Play to Generalize: Learning to Reason Through Game Play
- SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
- Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
- Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
- Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
- CoMemo: LVLMs Need Image Context with Image Memory
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
- MiMo-VL Technical Report
- Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
- PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
- GuessBench: Sensemaking Multimodal Creativity in the Wild
- HueManity: Probing Fine-Grained Visual Perception in MLLMs
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Grounded Reinforcement Learning for Visual Reasoning
- OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
- VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- R-Genie: Reasoning-Guided Generative Image Editing
- CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
- SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
- Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
- Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- SITE: towards Spatial Intelligence Thorough Evaluation
- Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
- Letting the neural code speak: Automated characterization of monkey visual neurons through human language
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Reinforced Attention Learning
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- BabyVision: Visual Reasoning Beyond Language
- Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
- GEB-Bench: Abstract Structures Told in Many Voices
- GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
- MIEB: Massive Image Embedding Benchmark
- Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Kimi-VL Technical Report
Discussions
Related