DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
2022/03/07 by Hao Zhang, Zhang, Hao, Feng Li +13 · 172 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2203.03605
openalex publication_date 2022/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
We present DINO (DETR with Improved deNoising anchOr boxes), a state-of-the-art end-to-end object detector. % in this paper. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a mixed query selection method for anchor initialization, and a look forward twice scheme for box prediction. DINO achieves 49.4AP in 12 epochs and 51.3AP in 24 epochs on COCO with a ResNet-50 backbone and multi-scale features, yielding a significant improvement of +6.0AP and +2.7AP, respectively, compared to DN-DETR, the previous best DETR-like model. DINO scales well in both model size and data size. Without bells and whistles, after pre-training on the Objects365 dataset with a SwinL backbone, DINO obtains the best results on both COCO val2017 (63.2AP) and test-dev (\textbf63.3AP). Compared to other models on the leaderboard, DINO significantly reduces its model size and pre-training data size while achieving better results. Our code will be available at \urlhttps://github.com/IDEACVR/DINO.
Cited by
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- DreamOmni3: Scribble-based Editing and Generation
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
- Small-Pollinator Detection in Cluttered Field Video
- CellMamba: Adaptive Mamba for Accurate and Efficient Cell Detection
- IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
- LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
- StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
- DenseBEV: Transforming BEV Grid Cells into 3D Objects
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders
- Learning to Generate Cross-Task Unexploitable Examples
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Optimal transport unlocks end-to-end learning for single-molecule localization
- MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial Images
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
- ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery
- Uni-Hema: Unified Model for Digital Hematopathology
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
- HybriDLA: Hybrid Generation for Document Layout Analysis
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting
- Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
- Benchmarking Table Extraction from Heterogeneous Scientific Extraction Documents
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- ProtoAnomalyNCD: Prototype Learning for Multi-class Novel Anomaly Discovery in Industrial Scenarios
- Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated Descriptors
- QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
- Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- BeyondFacial: Identity-Preserving Personalized Generation Beyond Facial Close-ups
- NP-LoRA: Null Space Projection Unifies Subject and Style in LoRA Fusion
- Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Exploring the Underwater World Segmentation without Extra Training
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
- Text to Sketch Generation with Multi-Styles
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- Test-Time Adaptive Object Detection with Foundation Model
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
- DAMap: Distance-aware MapNet for High Quality HD Map Construction
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis
- Towards 3D Objectness Learning in an Open World
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- Detect Anything via Next Point Prediction
- CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
- Source-Free Object Detection with Detection Transformer
- Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
- TARO: Toward Semantically Rich Open-World Object Detection
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Referring Expression Comprehension for Small Objects
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
- Stratum corneum nanotexture feature detection using deep learning and spatial analysis: a noninvasive tool for skin barrier assessment
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Sim-DETR: Unlock DETR for Temporal Sentence Grounding
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Large Material Gaussian Model for Relightable 3D Generation
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- RiO-DETR: DETR for Real-time Oriented Object Detection
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- Knowledge Transfer from Interaction Learning
- VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
- TASAM: Terrain-and-Aware Segment Anything Model for Temporal-Scale Remote Sensing Segmentation
- Region-Aware Deformable Convolutions
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications
- BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection
- CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
- Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
- SpectMamba: Integrating Frequency and State Space Models for Enhanced Medical Image Detection
- Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
- Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- DQEN: Dual Query Enhancement Network for DETR-based HOI Detection
- Robust and Label-Efficient Deep Waste Detection
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- You Only Pose Once: A Minimalist's Detection Transformer for Monocular RGB Category-level 9D Multi-Object Pose Estimation
- AnchorSync: Global Consistency Optimization for Long Video Editing
- ViT-FIQA: Assessing Face Image Quality using Vision Transformers
- TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
- Next Visual Granularity Generation
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
- Text-guided Visual Prompt DINO for Generic Segmentation
- UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
- Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a User-Centric Optimization Space
- Contextual Object Detection with Multimodal Large Language Models
- Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications
- Benchmarking pig detection and tracking under diverse and challenging conditions
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
- Length Matters: Length-Aware Transformer for Temporal Sentence Grounding
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- MedCAL-Bench: A Comprehensive Benchmark on Cold-Start Active Learning with Foundation Models for Medical Image Analysis
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Subject or Style: Adaptive and Training-Free Mixture of LoRAs
- Representation Shift: Unifying Token Compression with FlashAttention
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- Automated Detection of Antarctic Benthic Organisms in High-Resolution In Situ Imagery to Aid Biodiversity Monitoring
- AnimalClue: Recognizing Animals by their Traces
- Local2Global query Alignment for Video Instance Segmentation
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and beyond
- Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
- Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection
- Real-Time Fusion of Visual and Chart Data for Enhanced Maritime Vision
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- SS-DC: Spatial-Spectral Decoupling and Coupling Across Visible-Infrared Gap for Domain Adaptive Object Detection
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
- A document is worth a structured record: Principled inductive bias design for document recognition
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
Related