Object Recognition Datasets and Challenges: A Review
2025/07/30 by Salari, Aria, Djavadifar, Abtin, Liu, Xiangrui +1
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2507.22361
Abstract
Object recognition is among the fundamental tasks in the computer vision applications, paving the path for all other image understanding operations. In every stage of progress in object recognition research, efforts have been made to collect and annotate new datasets to match the capacity of the state-of-the-art algorithms. In recent years, the importance of the size and quality of datasets has been intensified as the utility of the emerging deep network techniques heavily relies on training data. Furthermore, datasets lay a fair benchmarking means for competitions and have proved instrumental to the advancements of object recognition research by providing quantifiable benchmarks for the developed models. Taking a closer look at the characteristics of commonly-used public datasets seems to be an important first step for data-driven and machine learning researchers. In this survey, we provide a detailed analysis of datasets in the highly investigated object recognition areas. More than 160 datasets have been scrutinized through statistics and descriptions. Additionally, we present an overview of the prominent object recognition benchmarks and competitions, along with a description of the metrics widely adopted for evaluation purposes in the computer vision community. All introduced datasets and challenges can be found online at github.com/AbtinDjavadifar/ORDC.
Citations
- Mutual Graph Learning for Camouflaged Object Detection
- AU-AIR: A Multi-modal Unmanned Aerial Vehicle Dataset for Low Altitude\n Traffic Surveillance
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections
- Deep Semantic Segmentation of Natural and Medical Images: A Review
- INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
- A*3D Dataset: Towards Autonomous Driving in Challenging Environments
- AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces
- Overview of LifeCLEF Plant Identification task 2019: diving into data deficient tropical countries
- CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation
- Argoverse: 3D Tracking and Forecasting with Rich Maps
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- D2-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios
- Survey on semantic segmentation using deep learning techniques
- The KiTS19 Challenge Data: 300 Kidney Tumor Cases with Clinical Context, CT Semantic Segmentations, and Surgical Outcomes
- TensorMask: A Foundation for Dense Object Segmentation
- nuScenes: A multimodal dataset for autonomous driving
- DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
- CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and\n Expert Comparison
- Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing
- The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories\n on German Highways for Validation of Highly Automated Driving Systems
- Deep Learning for Generic Object Detection: A Survey
- Deep Learning for Generic Object Detection: A Survey
- YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
- Object Detection with Deep Learning: A Review
- ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations
- Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition
- YOLOv3: An Incremental Improvement
- xView: Objects in Context in Overhead Imagery
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- Functional Map of the World
- A critical evaluation of the Next Generation Simulation (NGSIM) vehicle trajectory dataset
- VGGFace2: A dataset for recognising faces across pose and age
- Fast YOLO: A Fast You Only Look Once System for Real-time Embedded\n Object Detection in Video
- Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Level Playing Field for Million Scale Face Recognition
- A Review on Deep Learning Techniques Applied to Semantic Segmentation
- Mask R-CNN
- CityPersons: A Diverse Dataset for Pedestrian Detection
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- 1 year, 1000 km: The Oxford RobotCar dataset
- COCO-Stuff: Thing and Stuff Classes in Context
- Feature Pyramid Networks for Object Detection
- TorontoCity: Seeing the World with a Million Eyes
- UMDFaces: An Annotated Face Dataset for Training Deep Networks
- Xception: Deep Learning with Depthwise Separable Convolutions
- Places: An Image Database for Deep Scene Understanding
- A Large Contextual Dataset for Classification, Detection and Counting of\n Cars with Deep Learning
- AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
- MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic\n Videos
- EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos
- Deep Residual Learning for Image Recognition
- The MegaFace Benchmark: 1 Million Faces for Recognition at Scale
- You Only Look Once: Unified, Real-Time Object Detection
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Cross-domain Image Retrieval with a Dual Attribute-aware Ranking Network
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Fast R-CNN
- Visual Saliency Based on Multiscale Deep Features
- YFCC100M
- DeepID3: Face Recognition with Very Deep Neural Networks
- Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?
- Learning Face Representation from Scratch
- Deep Learning Face Attributes in the Wild
- ImageNet Large Scale Visual Recognition Challenge
- ImageNet Large Scale Visual Recognition Challenge
- Hierarchical Saliency Detection on Extended CSSD
- Deep Learning Face Representation by Joint Identification-Verification
- The Secrets of Salient Object Segmentation
- Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Microsoft COCO: Common Objects in Context
- Rich feature hierarchies for accurate object detection and semantic\n segmentation
- Vision meets robotics: The KITTI dataset
- ImageNet classification with deep convolutional neural networks
- The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A Completed Reference Database of Lung Nodules on CT Scans
- LabelMe: A Database and Web-Based Tool for Image Annotation
- A Fast Learning Algorithm for Deep Belief Nets
- Backpropagation Applied to Handwritten Zip Code Recognition
Related