Grounded Language-Image Pre-training
2021/12/07 by Li, Liunian Harold, Zhang, Pengchuan, Zhang, Haotian +9 · 126 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2112.03857
Abstract
This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representation semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1) When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines. 2) After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3) When transferred to 13 downstream object detection tasks, a 1-shot GLIP rivals with a fully-supervised Dynamic Head. Code is released at https://github.com/microsoft/GLIP.
Cited by
- SonoVision: A Computer Vision Approach for Helping Visually Challenged Individuals Locate Objects with the Help of Sound Cues
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation
- Particulate: Feed-Forward 3D Object Articulation
- SuperCLIP: CLIP with Simple Classification Supervision
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus
- VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Qwen3.5-Omni Technical Report
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- OpenBox: Annotate Any Bounding Boxes in 3D
- SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
- S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- PowerCLIP: Powerset Alignment for Contrastive Pre-Training
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- Qwen3-VL Technical Report
- RLM: A Vision-Language Model Approach for Radar Scene Understanding
- MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
- Intelligent Image Search Algorithms Fusing Visual Large Models
- Dual-Granularity Semantic Prompting for Language Guidance Infrared Small Target Detection
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
- Autonomous Surface Selection For Manipulator-Based UV Disinfection In Hospitals Using Foundation Models
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
- Trust in Vision-Language Models: Insights from a Participatory User Workshop
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- ZeroDexGrasp: Zero-Shot Task-Oriented Dexterous Grasp Synthesis with Prompt-Based Multi-Stage Semantic Reasoning
- SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation
- Binary Verification for Zero-Shot Vision
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- Test-Time Adaptive Object Detection with Foundation Model
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- [De|Re]constructing VLMs' Reasoning in Counting
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- The Mechanistic Emergence of Symbol Grounding in Language Models
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Detect Anything via Next Point Prediction
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Vision Language Models: A Survey of 26K Papers
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Referring Expression Comprehension for Small Objects
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
- Solar PV Installation Potential Assessment on Building Facades Based on Vision and Language Foundation Models
- GroundSight: Augmenting Vision-Language Models with Grounding Information and De-hallucination
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
- Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- Vision Language Models Cannot Plan, but Can They Formalize?
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- Synthetic Protein-Ligand Complex Generation for Deep Molecular Docking
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- Towards Understanding Visual Grounding in Visual Language Models
- Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
- Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
- X-Part: high fidelity and structure coherent shape decomposition
- P3-SAM: Native 3D Part Segmentation
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Towards Open World Detection: A Survey
- OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
- Semantic-Aware Ship Detection with Vision-Language Integration
- JVLGS: Joint Vision-Language Gas Leak Segmentation
- Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
- Agentic Design Review System
- Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- Text-guided Visual Prompt DINO for Generic Segmentation
- Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
- Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
Related