From Recognition to Cognition: Visual Commonsense Reasoning
2018/11/27 by Zellers, Rowan, Bisk, Yonatan, Farhadi, Ali +1 · 48 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.1811.10830
Abstract
Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people's actions, goals, and mental states. While this task is easy for humans, it is tremendously difficult for today's vision systems, requiring higher-order cognition and commonsense reasoning about the world. We formalize this task as Visual Commonsense Reasoning. Given a challenging question about an image, a machine must answer correctly and then provide a rationale justifying its answer. Next, we introduce a new dataset, VCR, consisting of 290k multiple choice QA problems derived from 110k movie scenes. The key recipe for generating non-trivial and high-quality problems at scale is Adversarial Matching, a new approach to transform rich annotations into multiple choice questions with minimal bias. Experimental results show that while humans find VCR easy (over 90% accuracy), state-of-the-art vision models struggle (~45%). To move towards cognition-level understanding, we present a new reasoning engine, Recognition to Cognition Networks (R2C), that models the necessary layered inferences for grounding, contextualization, and reasoning. R2C helps narrow the gap between humans and machines (~65%); still, the challenge is far from solved, and we provide analysis that suggests avenues for future work.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles
- Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- CauSight: Learning to Supersense for Visual Causal Discovery
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- Understanding Task Transfer in Vision-Language Models
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- Draft and Refine with Visual Experts
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
- Spot The Ball: A Benchmark for Visual Social Inference
- CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Causal Debiasing for Visual Commonsense Reasoning
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- Self-Improvement in Multimodal Large Language Models: A Survey
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VIVA+: Human-Centered Situational Decision-Making
- Vision-Grounded Machine Interpreting: Improving the Translation Process through Visual Cues
- Multilingual Vision-Language Models, A Survey
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
- Towards Understanding Visual Grounding in Visual Language Models
- Can VLMs Recall Factual Associations From Visual References?
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
Related