YOLO-World: Real-Time Open-Vocabulary Object Detection
2024/01/30 by Tianheng Cheng, Cheng, Tianheng, Lin Song +9 · 12 voices · 130 citations
Computer Science · #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2401.17270
Abstract
The You Only Look Once (YOLO) series of detectors have established themselves as efficient and practical tools. However, their reliance on predefined and trained object categories limits their applicability in open scenarios. Addressing this limitation, we introduce YOLO-World, an innovative approach that enhances YOLO with open-vocabulary detection capabilities through vision-language modeling and pre-training on large-scale datasets. Specifically, we propose a new Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN) and region-text contrastive loss to facilitate the interaction between visual and linguistic information. Our method excels in detecting a wide range of objects in a zero-shot manner with high efficiency. On the challenging LVIS dataset, YOLO-World achieves 35.4 AP with 52.0 FPS on V100, which outperforms many state-of-the-art methods in terms of both accuracy and speed. Furthermore, the fine-tuned YOLO-World achieves remarkable performance on several downstream tasks, including object detection and open-vocabulary instance segmentation.
Cited by
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
- VL-Nav: Neuro-Symbolic Reasoning-based Vision-Language Navigation
- UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation
- SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
- Attention from Above: A Multimodal Model for Drone-Based Object Localization
- Cognitive-YOLO: LLM-Driven Architecture Synthesis from First Principles of Data for Object Detection
- Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
- Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
- IoUPD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
- REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
- Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG
- IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
- Semantic Evidence Regulation via Relational Bias for Zero-Shot Object Navigation
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Retrieving Objects from 3D Scenes with Box-Guided Open-Vocabulary Instance Segmentation
- Object-Centric Framework for Video Moment Retrieval
- Auto-Vocabulary 3D Object Detection
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
- D2GSLAM: 4D Dynamic Gaussian Splatting SLAM
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
- MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- BlinkBud: Detecting Hazards from Behind via Sampled Monocular 3D Detection on a Single Earbud
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- Video Generation Models Are Good Latent Reward Models
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
- GSpyNetTreeS: a machine learning solution for glitch localization in time and frequency
- UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
- ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
- Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- Binary Verification for Zero-Shot Vision
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Generating Accurate and Detailed Captions for High-Resolution Images
- DIV-Nav: Open-Vocabulary Spatial Relationships for Multi-Object Navigation
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
- Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
- Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
- World-in-World: World Models in a Closed-Loop World
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- Detect Anything via Next Point Prediction
- Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
- SegMASt3R: Geometry Grounded Segment Matching
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
- Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
- From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
- PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
- Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Guiding Audio Editing with Audio Language Model
- Real-Time Indoor Object SLAM with LLM-Enhanced Priors
- SLAM-Free Visual Navigation with Hierarchical Vision-Language Perception and Coarse-to-Fine Semantic Topological Planning
- Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
- StereoFoley: Object-Aware Stereo Audio Generation from Video
- MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
- Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy
- EZREAL: Enhancing Zero-Shot Outdoor Robot Navigation toward Distant Targets under Varying Visibility
- MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
- Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions
- From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
- Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
- Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
- Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
- Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s
- OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection
- SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing
- Towards Open World Detection: A Survey
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
- OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
- Spotlighter: Revisiting Prompt Tuning from a Representative Mining View
- Robust and Label-Efficient Deep Waste Detection
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
- Text-guided Visual Prompt DINO for Generic Segmentation
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Open-Attribute Person Retrieval: Finding People Through Distinctive and Novel Attributes
- ODOV: Towards Open-Domain Open-Vocabulary Object Detection
- A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
Discussions
- YOLO-World: Real-Time Open-Vocabulary Object Detection [hn, 148 points, 52 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection [bsky, 1 points, 0 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3lqj5zrzrwkb2 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection #HackerNews https://arxiv.org/abs/2401.17270 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection https://arxiv.org/abs/2401.17270 https://news.ycombinator.com/item?id=44146858 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection https://arxiv.org/abs/2401.17270 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection https://arxiv.org/abs/2401.17270 (https://news.ycombinator.com/item?id=44146858) [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: YOLO-World: Real-Time Open-Vocabulary Object Detection [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2401.17270 YOLO-Worldは、リアルタイムで動作するopen-vocabularyの物体検出モデルです。 従来のYOLOの限界である事前定義された物体カテゴリへの依存を克服するために開発されました。 RepVL-PANという新しいネットワークと領域と言語のコントラスト損失により、視覚と言語情報の連携を強化しています。 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection view on hacker news [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection https://arxiv.org/abs/2401.17270 [bsky, 0 points, 0 comments]
- YOLO-World: Real-Time Open-Vocabulary Object Detection https://arxiv.org/abs/2401.17270 (https://news.ycombinator.com/item?id=44146858) [bsky, 0 points, 0 comments]
Related