LVIS: A Dataset for Large Vocabulary Instance Segmentation
2019/08/08 by Agrim Gupta, Piotr Dollár, Gupta, Agrim +3 · 164 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Multimodal Machine Learning Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.1908.03195
Extension of the CVPR'19 paper describing release v0.5, the LVIS Challenge, and baseline results
arxiv created 2019/09/15 · arxiv updated 2019/09/17
Abstract
Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks. In this work, we introduce LVIS (pronounced `el-vis'): a new dataset for Large Vocabulary Instance Segmentation. We plan to collect ~2 million high-quality instance segmentation masks for over 1000 entry-level object categories in 164k images. Due to the Zipfian distribution of categories in natural images, LVIS naturally has a long tail of categories with few training samples. Given that state-of-the-art deep learning methods for object detection perform poorly in the low-sample regime, we believe that our dataset poses an important and exciting new scientific challenge. LVIS is available at http://www.lvisdataset.org.
Citations
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
- Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning
- StAR: Segment Anything Reasoner
- X-ray Insights Unleashed: Pioneering the Enhancement of Multi-Label Long-Tail Data
- FlowDet: Unifying Object Detection and Generative Transport Flows
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- Omni-Referring Image Segmentation
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- MSG-Loc: Multi-Label Likelihood-based Semantic Graph Matching for Object-Level Global Localization
- Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
- FOM-Nav: Frontier-Object Maps for Object Goal Navigation
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- NNGPT: Rethinking AutoML with Large Language Models
- Understanding Task Transfer in Vision-Language Models
- State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
- iFlyBot-VLM Technical Report
- OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- TRACE: Textual Reasoning for Affordance Coordinate Extraction
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
- Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- Generative Universal Verifier as Multimodal Meta-Reasoner
- Detect Anything via Next Point Prediction
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Unified Open-World Segmentation with Multi-Modal Prompts
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- TARO: Toward Semantically Rich Open-World Object Detection
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
- Few-Shot Pattern Detection via Template Matching and Regression
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
- Augment to Segment: Tackling Pixel-Level Imbalance in Wheat Disease and Pest Segmentation
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
- Towards Open World Detection: A Survey
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
- Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Robust and Label-Efficient Deep Waste Detection
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
- A Guide for Manual Annotation of Scientific Imagery: How to Prepare for Large Projects
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- MolmoAct: Action Reasoning Models that can Reason in Space
- Text-guided Visual Prompt DINO for Generic Segmentation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Open Scene Graphs for Open-World Object-Goal Navigation
- X-SAM: From Segment Anything to Any Segmentation
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Enhancing Object Discovery for Unsupervised Instance Segmentation and Object Detection
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
- ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
- Discovering and using Spelke segments
- Evaluating the Efficacy of Large Language Models for Generating Fine-Grained Visual Privacy Policies in Homes
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Object Recognition Datasets and Challenges: A Review
- From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
- Test-Time Canonicalization by Foundation Models for Robust Perception
- BlueGlass: A Framework for Composite AI Safety
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
- CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Compress Any Segment Anything Model (SAM)
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- Inter- and Intra-image Refinement for Few Shot Segmentation
- Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
- Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
- Zero-shot Inexact CAD Model Alignment from a Single Image
- Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
- No time to train! Training-Free Reference-Based Instance Segmentation
- Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule
- Visual Textualization for Image Prompted Object Detection
- Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- Diffusion-Based Image Augmentation for Semantic Segmentation in Outdoor Robotics
- FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
- Synthetic Visual Genome
- PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
- FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
- Refer to Any Segmentation Mask Group With Vision-Language Prompts
- Gen-n-Val: Agentic Image Data Generation and Validation
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- Object-Shot Enhanced Grounding Network for Egocentric Video
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- Auto-Labeling Data for Object Detection
- GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning
- Common Inpainted Objects In-N-Out of Context
- Test-time Vocabulary Adaptation for Language-driven Object Detection
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
- ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
- Open-Det: An Efficient Learning Framework for Open-Ended Detection
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
- FruitNeRF++: A Generalized Multi-Fruit Counting Method Utilizing Contrastive Learning and Neural Radiance Fields
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- Reasoning Segmentation for Images and Videos: A Survey
- Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
- InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
- LTDA-Drive: LLMs-guided Generative Models based Long-tail Data Augmentation for Autonomous Driving
- Unlocking the Power of SAM 2 for Few-Shot Segmentation
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- AoP-SAM: Automation of Prompts for Efficient Segmentation
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
- FG-CLIP: Fine-Grained Visual and Textual Alignment
- Steerable Visual Representations
- Vision as Unified Multimodal Generation
- Vision Transformers Need More Than Registers
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- T2ID-CAS: Diffusion Model and Class Aware Sampling to Mitigate Class Imbalance in Neck Ultrasound Anatomical Landmark Detection
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
- Revisiting Data Auditing in Large Vision-Language Models
- Improving Open-World Object Localization by Discovering Background
- DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
- Describe Anything: Detailed Localized Image and Video Captioning
- LIFT+: Lightweight Fine-Tuning for Long-Tail Learning
- Universal Concept Disruption for SAM3 Image Segmentation
- HoloCount: A Holistic Visual Counting Benchmark for MLLMs
- Balancing Stability and Plasticity in Pretrained Detector: A Dual-Path Framework for Incremental Object Detection
- Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
- Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
Related