Modeling Context in Referring Expressions
2016/07/31 by Yu, Licheng, Poirson, Patrick, Yang, Shan +2 · 115 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.1608.00272
Abstract
Humans refer to objects in their environments all the time, especially in dialogue with other people. We explore generating and comprehending natural language referring expressions for objects in images. In particular, we focus on incorporating better measures of visual context into referring expression models and find that visual comparison to other objects within an image helps improve performance significantly. We also develop methods to tie the language generation process together, so that we generate expressions for all objects of a particular category jointly. Evaluation on three recent datasets - RefCOCO, RefCOCO+, and RefCOCOg, shows the advantages of our methods for both referring expression generation and comprehension.
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- An LMM for Precisely Grounding Elements in Documents
- StAR: Segment Anything Reasoner
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Benchmarking the Generality of Vision-Language-Action Models
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Grounding Everything in Tokens for Multimodal Large Language Models
- Generalized Referring Expression Segmentation on Aerial Photos
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- Omni-Referring Image Segmentation
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- SAM3-I: Segment Anything with Instructions
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- OneThinker: All-in-one Reasoning Model for Image and Video
- Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
- Generalized Medical Phrase Grounding
- Describe Anything Anywhere At Any Moment
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
- Direct Visual Grounding by Directing Attention of Visual Tokens
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- Binary Verification for Zero-Shot Vision
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- iFlyBot-VLM Technical Report
- Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- A Simple and Better Baseline for Visual Grounding
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Vision Language Models: A Survey of 26K Papers
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- Referring Expression Comprehension for Small Objects
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- From Psycholinguistics to Computer Vision. A Comprehensive Review of Object Naming Data and Studies
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Mitigating Query Selection Bias in Referring Video Object Segmentation
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
- Towards Understanding Visual Grounding in Visual Language Models
- Visual Grounding from Event Cameras
- XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- See No Evil: Adversarial Attacks Against Linguistic-Visual Association in Referring Multi-Object Tracking Systems
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
- RynnEC: Bringing MLLMs into Embodied World
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Latent Expression Generation for Referring Image Segmentation and Grounding
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
- Multimodal Referring Segmentation: A Survey
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
Related