Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection
2020/06/01 by Shifeng Zhang, Cheng Chi, Yongqiang Yao +2 · 94 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
paper · doi:10.1109/cvpr42600.2020.00978
openalex publication_date 2020/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Object detection has been dominated by anchor-based detectors for several years. Recently, anchor-free detectors have become popular due to the proposal of FPN and Focal Loss. In this paper, we first point out that the essential difference between anchor-based and anchor-free detection is actually how to define positive and negative training samples, which leads to the performance gap between them. If they adopt the same definition of positive and negative samples during training, there is no obvious difference in the final performance, no matter regressing from a box or a point. This shows that how to select positive and negative training samples is important for current object detectors. Then, we propose an Adaptive Training Sample Selection (ATSS) to automatically select positive and negative samples according to statistical characteristics of object. It significantly improves the performance of anchor-based and anchor-free detectors and bridges the gap between them. Finally, we discuss the necessity of tiling multiple anchors per location on the image to detect objects. Extensive experiments conducted on MS COCO support our aforementioned analysis and conclusions. With the newly introduced ATSS, we improve state-of-the-art detectors by a large margin to 50.7% AP without introducing any overhead. The code is available at https://github.com/sfzhang15/ATSS.
Cited by
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Real-Time Visual Object Tracking via Few-Shot Learning
- You Only Look One-level Feature
- Disentangle Your Dense Object Detector
- Understanding Mobile GUI: from Pixel-Words to Screen-Sentences
- SRF-GAN: Super-Resolved Feature GAN for Multi-Scale Representation
- VarifocalNet: An IoU-aware Dense Object Detector
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- Detection and Tracking Meet Drones Challenge
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- Implicit Feature Pyramid Network for Object Detection
- Oriented RepPoints for Aerial Object Detection
- Forestpest-YOLO: A High-Performance Detection Framework for Small Forestry Pests
- Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
- Calibrated and Resource-Aware Super-Resolution for Reliable Driver Behavior Analysis
- Sample and Computation Redistribution for Efficient Face Detection
- Representation Sharing for Fast Object Detector Search and Beyond
- Visual Detector Compression via Location-Aware Discriminant Analysis
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Region-Aware Deformable Convolutions
- MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
- ARS-DETR: Aspect Ratio-Sensitive Detection Transformer for Aerial Oriented Object Detection
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- CEM-FBGTinyDet: Context-Enhanced Foreground Balance with Gradient Tuning for tiny Objects
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- FCOS: A simple and strong anchor-free object detector
- Revisiting the Loss Weight Adjustment in Object Detection
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- Probabilistic two-stage detection
- YOLOX: Exceeding YOLO Series in 2021
- PVT v2: Improved baselines with pyramid vision transformer
- COXNet: Cross-Layer Fusion with Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
- Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
- Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search
- Scale-aware Automatic Augmentation for Object Detection
- Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
- PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
- Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild
- RG-YOLO: multi-scale feature learning for underwater target detection
- X-safe: an X-ray security detection method based on incremental Kernel aggregation, hierarchical co-optimization and task-aligned labeling
- Anchor-free 3D Single Stage Detector with Mask-Guided Attention for Point Cloud
- RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- HR-RCNN: Hierarchical Relational Reasoning for Object Detection
- Learning Oriented Remote Sensing Object Detection via Naive Geometric Computing
- Drone-based RGBT tiny person detection
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Local Metrics for Multi-Object Tracking
- Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
- High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
- TOOD: Task-aligned One-stage Object Detection
- MiniVLM: A Smaller and Faster Vision-Language Model
- Variational Pedestrian Detection
- USIS16K: High-Quality Dataset for Underwater Salient Instance Segmentation
- GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel Attention
- FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection
- SFFN-YOLO for small object detection in aerial images
- VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
- RAPiD: Rotation-Aware People Detection in Overhead Fisheye Images
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector
- FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
- Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
- Dynamic Head: Unifying Object Detection Heads with Attentions
- Discriminative Semantic Feature Pyramid Network with Guided Anchoring for Logo Detection
- Modulating Localization and Classification for Harmonized Object Detection
- Structure Information is the Key: Self-Attention RoI Feature Extractor in 3D Object Detection
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- Augmenting Proposals by the Detector Itself
- AuxDet: Auxiliary Metadata Matters for Omni-Domain Infrared Small Target Detection
- DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation
- A Normalized Gaussian Wasserstein Distance for Tiny Object Detection
- DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object Detection
- M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection
- IQDet: Instance-wise Quality Distribution Sampling for Object Detection
- Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition
- A Simple Detector with Frame Dynamics is a Strong Tracker
- Localization Distillation for Dense Object Detection
- RMOPP: Robust Multi-Objective Post-Processing for Effective Object Detection
- Focal and efficient IOU loss for accurate bounding box regression
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- Purifying, Labeling, and Utilizing: A High-Quality Pipeline for Small Object Detection
- Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction
- More Clear, More Flexible, More Precise: A Comprehensive Oriented Object Detection benchmark for UAV
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- Wide-Depth-Range 6D Object Pose Estimation in Space
- Universal Lymph Node Detection in Multiparametric MRI with Selective Augmentation
- 1st Place Solution for ICDAR 2021 Competition on Mathematical Formula Detection
- LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D Detector
- Density-based Object Detection in Crowded Scenes
- Class Imbalance Correction for Improved Universal Lesion Detection and Tagging in CT
Related