Object Detection in 20 Years: A Survey
2019/05/13 by Zhengxia Zou, Zou, Zhengxia, Chen, Keyan +6 · 118 citations
Computer Science · #Advanced Neural Network Applications #Advanced Image and Video Retrieval Techniques #Currency Recognition and Detection
paper · pdf · doi:10.48550/arxiv.1905.05055
Abstract
Object detection, as of one the most fundamental and challenging problems in computer vision, has received great attention in recent years. Over the past two decades, we have seen a rapid technological evolution of object detection and its profound impact on the entire computer vision field. If we consider today's object detection technique as a revolution driven by deep learning, then back in the 1990s, we would see the ingenious thinking and long-term perspective design of early computer vision. This paper extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century's time (from the 1990s to 2022). A number of topics have been covered in this paper, including the milestone detectors in history, detection datasets, metrics, fundamental building blocks of the detection system, speed-up techniques, and the recent state-of-the-art detection methods.
Citations
Cited by
- Neuromorphic Object Detection: An In-Depth Study and Future Directions
- CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
- VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- Ideal Observer for Segmentation of Dead Leaves Images
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- When and how to automate image analysis for wildlife monitoring? Guidelines and lessons from a worked example of seabirds in a dynamic coastal environment
- Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
- HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition
- Person Recognition in Aerial Surveillance: A Decade Survey
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
- YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection
- Linear time small coresets for k-mean clustering of segments with applications
- DeepDefense: Robust Learning via Layer-Wise Gradient-Feature Alignment
- DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Parameterized Prompt for Incremental Object Detection
- All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Region-CAM: Towards Accurate Object Regions in Class Activation Maps for Weakly Supervised Learning Tasks
- One-Timestep is Enough: Achieving High-performance ANN-to-SNN Conversion via Scale-and-Fire Neurons
- Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- Extreme Amodal Face Detection
- Getting the Numbers Right\unicodex2014Modelling Multi-Class Object Counting in Dense and Varied Scenes
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- The Use of the Simplex Architecture to Enhance Safety in Deep-Learning-Powered Autonomous Systems
- RiO-DETR: DETR for Real-time Oriented Object Detection
- Generative data augmentation for biliary tract detection on intraoperative images
- Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching
- FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World
- GPU Temperature Simulation-Based Testing for In-Vehicle Deep Learning Frameworks
- Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
- Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- From Orthomosaics to Raw UAV Imagery: Enhancing Palm Detection and Crown-Center Localization
- RT-DETR++ for UAV Object Detection
- Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM
- Towards Open World Detection: A Survey
- OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- AI Driven Road Maintenance Inspection
- FLUID: A Fine-Grained Lightweight Urban Signalized-Intersection Dataset of Dense Conflict Trajectories
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
- Can Large Language Models Identify Materials from Radar Signals?
- Understanding the Risks of Asphalt Art to the Reliability of Vision-Based Perception Systems
- An Event-based Fast Intensity Reconstruction Scheme for UAV Real-time Perception
- Unleashing the power of disruptive and emerging technologies amid COVID-19: A detailed review
- Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
- Advances and Challenges in Deep Lip Reading
- AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- WiSE-OD: Benchmarking Robustness in Infrared Object Detection
- RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
- Unmasking Performance Gaps: A Comparative Study of Human Anonymization and Its Effects on Video Anomaly Detection
- Edge Intelligence with Spiking Neural Networks
- Empirical Upper Bound in Object Detection and More
- ShrinkBox: Backdoor Attack on Object Detection to Disrupt Collision Avoidance in Machine Learning-based Advanced Driver Assistance Systems
- Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
- BakuFlow: A Streamlining Semi-Automatic Label Generation Tool
- Adaptive Object Detection with ESRGAN-Enhanced Resolution & Faster R-CNN
- Measuring the Impact of Rotation Equivariance on Aerial Object Detection
- BlueGlass: A Framework for Composite AI Safety
- A document is worth a structured record: Principled inductive bias design for document recognition
- RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models
- Image recognition via Vietoris-Rips complex
- I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors
- Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios
- A Unified Framework for Stealthy Adversarial Generation via Latent Optimization and Transferability Enhancement
- Ovis-U1 Technical Report
- A Survey of Human-in-the-loop for Machine Learning
- On Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances and Million-AID
- Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
- ProARD: progressive adversarial robustness distillation: provide wide range of robust students
- TITAN: Query-Token based Domain Adaptive Adversarial Learning
- LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection
- Multi-target DoA Estimation with an Audio-visual Fusion Mechanism
- Generalizing vision-language models to novel domains: A comprehensive survey
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- FindMeIfYouCan: Bringing Open Set metrics to near , far and farther Out-of-Distribution Object Detection
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- UniDet-D: A Unified Dynamic Spectral Attention Model for Object Detection under Adverse Weathers
- Noise Modulation: Let Your Model Interpret Itself
- High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
- PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation
- Dynamic Head: Unifying Object Detection Heads with Attentions
- VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
- Vision-Based Assistive Technologies for People with Cerebral Visual Impairment: A Review and Focus Study
- Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
- Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling
- An Efficient Accelerator Design Methodology for Deformable Convolutional Networks
- Out of the Shadows: Exploring a Latent Space for Neural Network Verification
- EarthSynth: Generating Informative Earth Observation with Diffusion Models
- A Comparison for Anti-noise Robustness of Deep Learning Classification Methods on a Tiny Object Image Dataset: from Convolutional Neural Network to Visual Transformer and Performer
- Driving Datasets Literature Review
- Fine-grained spatial-temporal perception for gas leak segmentation
- Person detection and re-identification in open-world settings of retail stores and public spaces
- Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
- Training dataset generation for bridge game registration
- A Decade of You Only Look Once (YOLO) for Object Detection: A Review
- A Review of 3D Object Detection with Vision-Language Models
- Context Aware Grounded Teacher for Source Free Object Detection
- Harmony: A Unified Framework for Modality Incremental Learning
- ADA-Net: Attention-Guided Domain Adaptation Network with Contrastive Learning for Standing Dead Tree Segmentation Using Aerial Imagery
- Human Aligned Compression for Robust Models
- Deep Learning in Concealed Dense Prediction
- Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression Localization
- SO-DETR: Leveraging Dual-Domain Features and Knowledge Distillation for Small Object Detection
- Empirical Upper Bound, Error Diagnosis and Invariance Analysis of Modern Object Detectors
- SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic Segmentation
Related