Generation and Comprehension of Unambiguous Object Descriptions
2015/11/07 by Junhua Mao, Mao, Junhua, Jonathan Huang +9 · 100 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #I.2.10 #I.2.6 #I.2.7 #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.1511.02283
openalex publication_date 2015/11/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods that generate descriptions of objects without taking into account other potentially ambiguous objects in the scene. Our model is inspired by recent successes of deep learning methods for image captioning, but while image captioning is difficult to evaluate, our task allows for easy objective evaluation. We also present a new large-scale dataset for referring expressions, based on MS-COCO. We have released the dataset and a toolbox for visualization and evaluation, see https://github.com/mjhucla/GoogleRefexptoolbox
Citations
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Grounding Everything in Tokens for Multimodal Large Language Models
- Omni-Referring Image Segmentation
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- DIQ-H: Evaluating Hallucination Persistence in VLMs Under Temporal Visual Degradation
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- Qwen3-VL Technical Report
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- Fast Reasoning Segmentation for Images and Videos
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- Binary Verification for Zero-Shot Vision
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- LongCat-Flash-Omni Technical Report
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- Detect Anything via Next Point Prediction
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- Unified Open-World Segmentation with Multi-Modal Prompts
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Vision Language Models: A Survey of 26K Papers
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- Referring Expression Comprehension for Small Objects
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- GroundSight: Augmenting Vision-Language Models with Grounding Information and De-hallucination
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Mitigating Query Selection Bias in Referring Video Object Segmentation
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
- Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
- Towards Understanding Visual Grounding in Visual Language Models
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- RynnEC: Bringing MLLMs into Embodied World
- ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- Ovis2.5 Technical Report
- KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
- Latent Expression Generation for Referring Image Segmentation and Grounding
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- Multimodal Referring Segmentation: A Survey
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
- Object-centric Video Question Answering with Visual Grounding and Referring
Related