vix.ing · top · new · best · stats · spec

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

2022/03/07 by Hao Zhang, Zhang, Hao, Feng Li +13 · 172 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.2203.03605

openalex publication_date 2022/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

We present DINO (DETR with Improved deNoising anchOr boxes), a state-of-the-art end-to-end object detector. % in this paper. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a mixed query selection method for anchor initialization, and a look forward twice scheme for box prediction. DINO achieves 49.4AP in 12 epochs and 51.3AP in 24 epochs on COCO with a ResNet-50 backbone and multi-scale features, yielding a significant improvement of +6.0AP and +2.7AP, respectively, compared to DN-DETR, the previous best DETR-like model. DINO scales well in both model size and data size. Without bells and whistles, after pre-training on the Objects365 dataset with a SwinL backbone, DINO obtains the best results on both COCO val2017 (63.2AP) and test-dev (\textbf63.3AP). Compared to other models on the leaderboard, DINO significantly reduces its model size and pre-training data size while achieving better results. Our code will be available at \urlhttps://github.com/IDEACVR/DINO.

Cited by

Related