Towards Open World Detection: A Survey
2025/08/22 by Bulzan, Andrei-Stefan, Cernazanu-Glavan, Cosmin
#68T45 #A.1 #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2 #I.4
paper · doi:10.48550/arxiv.2508.16527
Abstract
For decades, Computer Vision has aimed at enabling machines to perceive the external world. Initial limitations led to the development of highly specialized niches. As success in each task accrued and research progressed, increasingly complex perception tasks emerged. This survey charts the convergence of these tasks and, in doing so, introduces Open World Detection (OWD), an umbrella term we propose to unify class-agnostic and generally applicable detection models in the vision domain. We start from the history of foundational vision subdomains and cover key concepts, methodologies and datasets making up today's state-of-the-art landscape. This traverses topics starting from early saliency detection, foreground/background separation, out of distribution detection and leading up to open world object detection, zero-shot detection and Vision Large Language Models (VLLMs). We explore the overlap between these subdomains, their increasing convergence, and their potential to unify into a singular domain in the future, perception.
Citations
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- Attention Prompting on Image for Large Vision-Language Models
- Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
- SAM 2: Segment Anything in Images and Videos
- PaliGemma: A versatile 3B VLM for transfer
- A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
- Unveiling Encoder-Free Vision-Language Models
- DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
- Matryoshka Multimodal Models
- Streaming Long Video Understanding with Large Language Models
- Foundation Models for Video Understanding: A Survey
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- When Do We Not Need Larger Vision Models?
- Open-World Semantic Segmentation Including Class Similarity
- BSDP: Brain-inspired Streaming Dual-level Perturbations for Online Open World Object Detection
- YOLO-World: Real-Time Open-Vocabulary Object Detection
- MM-LLMs: Recent Advances in MultiModal Large Language Models
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- VCoder: Versatile Vision Encoders for Multimodal Large Language Models
- Interfacing Foundation Models' Embeddings
- ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- SEMPART: Self-supervised Multi-resolution Partitioning of Image Semantics
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future
- MMBench: Is Your Multi-modal Model an All-around Player?
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- Evaluating Object Hallucination in Large Vision-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Segment Everything Everywhere All at Once
- Micrograph segmentations for DDEVD
- Segment Anything
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- LLaMA: Open and Efficient Foundation Language Models
- CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection
- Unleashing the Power of Visual Prompting At the Pixel Level
- Generalized Decoding for Pixel, Image, and Language
- Unsupervised Object Localization: Observing the Background to Discover Objects
- LAION-5B: An open large-scale dataset for training next generation image-text models
- OpenOOD: Benchmarking Generalized Out-of-Distribution Detection
- Generalised Co-Salient Object Detection
- Open-world Semantic Segmentation via Contrasting and Clustering Vision-Language Embedding
- Matryoshka Representation Learning
- Flamingo: a Visual Language Model for Few-Shot Learning
- ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
- Out-of-Distribution Detection with Deep Nearest Neighbors
- Open-World Instance Segmentation: Exploiting Pseudo Ground Truth From Learned Pairwise Affinity
- PromptDet: Towards Open-vocabulary Detection using Uncurated Images
- Unknown-Aware Object Detection: Learning What You Don't Know from Videos in the Wild
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Revisiting Open World Object Detection
- Grounded Language-Image Pre-training
- Scaling Up Vision-Language Pre-training for Image Captioning
- RedCaps: web-curated image-text data created by the people, for the people
- Masked Autoencoders Are Scalable Vision Learners
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Generalized Out-of-Distribution Detection: A Survey
- Generalized Out-of-Distribution Detection: A Survey
- Specificity-preserving RGB-D Saliency Detection
- Deep Metric Learning for Open World Semantic Segmentation
- Exploring the Limits of Out-of-Distribution Detection
- Emerging Properties in Self-Supervised Vision Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation
- Towards Open World Object Detection
- Learning Transferable Visual Models From Natural Language Supervision
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize\n Long-Tail Visual Concepts
- Synthesizing the Unseen for Zero-shot Object Detection
- Energy-based Out-of-distribution Detection
- Toward unsupervised, multi-object discovery in large-scale image\n collections
- Object Segmentation Without Labels with Large-Scale Generative Models
- Multi-scale Interactive Network for Salient Object Detection
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
- A Simple Framework for Contrastive Learning of Visual Representations
- Dont Even Look Once: Synthesizing Features for Zero-Shot Detection
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- Object Detection in 20 Years: A Survey
- Object Detection in 20 Years: A Survey
- nuScenes: A multimodal dataset for autonomous driving
- PiCANet: Pixel-wise Contextual Attention Learning for Accurate Saliency Detection
- Object Detection with Deep Learning: A Review
- Rotation Equivariant CNNs for Digital Pathology
- Zero-Shot Object Detection
- EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
- Zero-Shot Learning -- A Comprehensive Evaluation of the Good, the Bad and the Ugly
- Attention Is All You Need
- Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary\n Visual Reasoning
- On the Role and the Importance of Features for Background Modeling and Foreground Detection
- A Deep Multi-Level Network for Saliency Prediction
- Semantic Understanding of Scenes through the ADE20K Dataset
- Semantic Understanding of Scenes Through the ADE20K Dataset
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Deep Residual Learning for Image Recognition
- You Only Look Once: Unified, Real-Time Object Detection
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for\n Richer Image-to-Sentence Models
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Fast R-CNN
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and Tell: A Neural Image Caption Generator
- Fully Convolutional Networks for Semantic Segmentation
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Describing Textures in the Wild
- Rich feature hierarchies for accurate object detection and semantic\n segmentation
- Challenges in Representation Learning: A report on three machine\n learning contests
- Fine-Grained Visual Classification of Aircraft
- ImageNet classification with deep convolutional neural networks
- Distinctive Image Features from Scale-Invariant Keypoints
- Estimating the Support of a High-Dimensional Distribution
- Long Short-Term Memory
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Gradient-based learning applied to document recognition
- A feature-integration theory of attention
- A Threshold Selection Method from Gray-Level Histograms
Related