From Recognition to Cognition: Visual Commonsense Reasoning
2018/11/27 by Rowan Zellers, Zellers, Rowan, Yonatan Bisk +5 · 82 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.1811.10830
CVPR 2019 oral. Project page at https://visualcommonsense.com
arxiv created 2019/03/26 · arxiv updated 2019/03/27
Abstract
Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people's actions, goals, and mental states. While this task is easy for humans, it is tremendously difficult for today's vision systems, requiring higher-order cognition and commonsense reasoning about the world. We formalize this task as Visual Commonsense Reasoning. Given a challenging question about an image, a machine must answer correctly and then provide a rationale justifying its answer. Next, we introduce a new dataset, VCR, consisting of 290k multiple choice QA problems derived from 110k movie scenes. The key recipe for generating non-trivial and high-quality problems at scale is Adversarial Matching, a new approach to transform rich annotations into multiple choice questions with minimal bias. Experimental results show that while humans find VCR easy (over 90% accuracy), state-of-the-art vision models struggle (~45%). To move towards cognition-level understanding, we present a new reasoning engine, Recognition to Cognition Networks (R2C), that models the necessary layered inferences for grounding, contextualization, and reasoning. R2C helps narrow the gap between humans and machines (~65%); still, the challenge is far from solved, and we provide analysis that suggests avenues for future work.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles
- Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- CauSight: Learning to Supersense for Visual Causal Discovery
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- Understanding Task Transfer in Vision-Language Models
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- Draft and Refine with Visual Experts
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
- Spot The Ball: A Benchmark for Visual Social Inference
- CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Causal Debiasing for Visual Commonsense Reasoning
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- Self-Improvement in Multimodal Large Language Models: A Survey
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VIVA+: Human-Centered Situational Decision-Making
- Vision-Grounded Machine Interpreting: Improving the Translation Process through Visual Cues
- Multilingual Vision-Language Models, A Survey
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
- Towards Understanding Visual Grounding in Visual Language Models
- Can VLMs Recall Factual Associations From Visual References?
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
- VIBE: Can a VLM Read the Room?
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
- Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Test-Time Consistency in Vision Language Models
- Synthetic Visual Genome
- CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
- Can Argus Judge Them All? Comparing VLMs Across Domains
- MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- Capturing Visualization Design Rationale
- Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
- On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
- Refer to Any Segmentation Mask Group With Vision-Language Prompts
- R3-VQA: "Read the Room" by Video Social Reasoning
- What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Multimodal Conversation Structure Understanding
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
- IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
- Task-Core Memory Management and Consolidation for Long-term Continual Learning
- Computational Reasoning of Large Language Models
- Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
- VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
- ChartQA-X: Generating Explanations for Visual Chart Reasoning
- Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis
- Evolved Hierarchical Masking for Self-Supervised Learning
- Impact of Language Guidance: A Reproducibility Study
- OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
Related