Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
2015/05/19 by Bryan A. Plummer, Plummer, Bryan A., Liwei Wang +10 · 250 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.1505.04870
openalex publication_date 2015/05/19 · arxiv created 2016/09/19 · arxiv updated 2016/09/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.
Citations
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- Null-LoRA: Low-Rank Adaptation on Null Space
- FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Benchmarking the Generality of Vision-Language-Action Models
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- Beyond Real Weights: Hypercomplex Representations for Stable Quantization
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Personalized Image Descriptions from Attention Sequences
- Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
- Technical Report on Text Dataset Distillation
- Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
- Hierarchical Semantic Alignment for Image Clustering
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
- Semantic-Aware Caching for Efficient Image Generation in Edge Computing
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
- Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
- Video Finetuning Improves Reasoning Between Frames
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Draft and Refine with Visual Experts
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
- CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- Surprisal reveals diversity gaps in image captioning and different scorers change the story
- Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
- From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
- Masked Diffusion Captioning for Visual Feature Learning
- Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- Progressive Multimodal Alignment for Continual Instruction Tuning
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
- Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
- ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- PHyCLIP: ℓ1-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Referring Expression Comprehension for Small Objects
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- ModernVBERT: Towards Smaller Visual Document Retrievers
- TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
- MuSLR: Multimodal Symbolic Logical Reasoning
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
- Deepfakes: we need to re-think the concept of "real" images
- Unifying Adversarially Robust Model Experts in Vision-Language Models
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Long Story Short: Disentangling Compositionality and Long-Caption Understanding in VLMs
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
- RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
- Efficient Multimodal Dataset Distillation via Generative Models
- MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Neural Text Generation with Artificial Negative Examples
- Evaluating Robustness of Vision-Language Models Under Noisy Conditions
- Towards Understanding Visual Grounding in Visual Language Models
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- Florence: A New Foundation Model for Computer Vision
- Effectively obtaining acoustic, visual and textual data from videos
- Semantic-guided LoRA Parameters Generation
- Towards Open World Detection: A Survey
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Understanding Data Influence with Differential Approximation
- Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering
- 7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
- Region-Level Context-Aware Multimodal Understanding
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
- ChartCap: Mitigating Hallucination of Dense Chart Captioning
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
- MultiSHAP: A Shapley-Based Framework for Explaining Cross-Modal Interactions in Multimodal AI Models
- Multimodal Referring Segmentation: A Survey
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
- Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- Trade-offs in Image Generation: How Do Different Dimensions Interact?
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- ZSE-Cap: A Zero-Shot Ensemble for Image Retrieval and Prompt-Guided Captioning
- Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation
- SPICE: Semantic Propositional Image Caption Evaluation
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- Visual Reasoning with Natural Language
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
- ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
- Change of Thought: Adaptive Test-Time Computation
- Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing
- Semantically Informed Salient Regions Guided Radiology Report Generation
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
- An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
- PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
- CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
- Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
- Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
- Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Synthetic Visual Genome
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- Generalizing vision-language models to novel domains: A comprehensive survey
- With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
- Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
- Control and Realism: Best of Both Worlds in Layout-to-Image without Training
- GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models
- CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
- On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
- Complexity of normalized stochastic first-order methods with momentum under heavy-tailed noise
- CoMemo: LVLMs Need Image Context with Image Memory
- GenIR: Generative Visual Feedback for Mental Image Retrieval
- DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
- Robust Anti-Backdoor Instruction Tuning in LVLMs
- DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
- DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation
- R2SM: Referring and Reasoning for Selective Masks
- Data Pruning by Information Maximization
- Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
- Benchmarking Foundation Models for Zero-Shot Biometric Tasks
- Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
- FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
- Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
- Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
- Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
- Reasoning Segmentation for Images and Videos: A Survey
- So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection
- TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
- EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
- Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
- Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- Learning Shared Representations from Unpaired Data
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
- Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
- PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model Merging
- GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding
- TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries
- Compositional Image-Text Matching and Retrieval by Grounding Entities
- Learning Cross-modal Context Graph for Visual Grounding
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
- Modeling Scientific Experiment Scenes: Dataset and Model
- Vision as Unified Multimodal Generation
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
- ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
- MATCHA: Matching Text via Contrastive Semantic Alignment
- AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection
- What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- URECA: Unique Region Caption Anything
- Towards Visual Text Grounding of Multimodal Large Language Model
- Decoupled Global-Local Alignment for Improving Compositional Understanding
- Describe Anything: Detailed Localized Image and Video Captioning
- Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
- Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
- ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
- Feature Importance-Aware Deep Joint Source-Channel Coding for Computationally Efficient and Adjustable Image Transmission
Related