YOLOv10: Real-Time End-to-End Object Detection
2024/05/23 by Wang, Ao, Chen, Hui, Liu, Lihao +4 · 106 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2405.14458
Abstract
Over the past years, YOLOs have emerged as the predominant paradigm in the field of real-time object detection owing to their effective balance between computational cost and detection performance. Researchers have explored the architectural designs, optimization objectives, data augmentation strategies, and others for YOLOs, achieving notable progress. However, the reliance on the non-maximum suppression (NMS) for post-processing hampers the end-to-end deployment of YOLOs and adversely impacts the inference latency. Besides, the design of various components in YOLOs lacks the comprehensive and thorough inspection, resulting in noticeable computational redundancy and limiting the model's capability. It renders the suboptimal efficiency, along with considerable potential for performance improvements. In this work, we aim to further advance the performance-efficiency boundary of YOLOs from both the post-processing and model architecture. To this end, we first present the consistent dual assignments for NMS-free training of YOLOs, which brings competitive performance and low inference latency simultaneously. Moreover, we introduce the holistic efficiency-accuracy driven model design strategy for YOLOs. We comprehensively optimize various components of YOLOs from both efficiency and accuracy perspectives, which greatly reduces the computational overhead and enhances the capability. The outcome of our effort is a new generation of YOLO series for real-time end-to-end object detection, dubbed YOLOv10. Extensive experiments show that YOLOv10 achieves state-of-the-art performance and efficiency across various model scales. For example, our YOLOv10-S is 1.8× faster than RT-DETR-R18 under the similar AP on COCO, meanwhile enjoying 2.8× smaller number of parameters and FLOPs. Compared with YOLOv9-C, YOLOv10-B has 46% less latency and 25% fewer parameters for the same performance.
Cited by
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
- Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
- Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- NavEYE: Vision–centered multi–sensor fusion–based situational awareness system for intelligent surface vehicles
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- Multi-temporal Adaptive Red-Green-Blue and Long-Wave Infrared Fusion for You Only Look Once-Based Landmine Detection from Unmanned Aerial Systems
- IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
- PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
- YolovN-CBi: A Lightweight and Efficient Architecture for Real-Time Detection of Small UAVs
- YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping
- TCLeaf-Net: a transformer-convolution framework with global-local attention for robust in-field lesion-level plant leaf disease detection
- Building Audio-Visual Digital Twins with Smartphones
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- AIRHILT: A Human-in-the-Loop Testbed for Multimodal Conflict Detection in Aviation
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- View-aware Cross-modal Distillation for Multi-view Action Recognition
- Bridging Granularity Gaps: Hierarchical Semantic Learning for Cross-domain Few-shot Segmentation
- Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
- DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
- Hand Held Multi-Object Tracking Dataset in American Football
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Garbage Vulnerable Point Monitoring using IoT and Computer Vision
- Deep learning-based object detection of offshore platforms on Sentinel-1 Imagery and the impact of synthetic training data
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- Sewer pipeline condition assessment and defect detection using computer vision
- All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
- Real-Time Threaded Houbara Detection and Segmentation for Wildlife Conservation using Mobile Platforms
- From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models
- SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
- Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models
- Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy
- Towards Intelligent Traffic Signaling in Dhaka City Based on Vehicle Detection and Congestion Optimization
- TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- AI-Enhanced Real-Time Wi-Fi Sensing Through Single Transceiver Pair
- GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- DEF-YOLO: Leveraging YOLO for Concealed Weapon Detection in Thermal Imagin
- MRS-YOLO Railroad Transmission Line Foreign Object Detection Based on Improved YOLO11 and Channel Pruning
- Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
- Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- In-Field Mapping of Grape Yield and Quality with Illumination-Invariant Deep Learning
- Investigating mixed traffic dynamics of pedestrians and non-motorized vehicles at urban intersections: Observation experiments and modelling
- TempoControl: Temporal Attention Guidance for Text-to-Video Models
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- Real-Time Object Detection Meets DINOv3
- Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection
- A Comparative Benchmark of Real-time Detectors for Blueberry Detection towards Precision Orchard Management
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
- Offshore oil and gas platform dynamics in the North Sea, Gulf of Mexico, and Persian Gulf: exploiting the Sentinel-1 archive
- Designing Practical Models for Isolated Word Visual Speech Recognition
- Real-Time Fish Detection in Indonesian Marine Ecosystems Using Lightweight YOLOv10-nano Architecture
- CETUS: Causal Event-Driven Temporal Modeling With Unified Variable-Rate Scheduling
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- Cott-ADNet: Lightweight Real-Time Cotton Boll and Flower Detection Under Field Conditions
- Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
- Scensory: Automated Real-Time Fungal Identification and Spatial Mapping
- Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- Artificial intelligence and companion animals: Perspectives on digital healthcare for dogs, cats, and pet ownership
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- STAGNet: A Spatio-Temporal Graph and LSTM Framework for Accident Anticipation
- AquaFeat: A Features-Based Image Enhancement Model for Underwater Object Detection
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets
- TrackOR: Towards Personalized Intelligent Operating Rooms Through Robust Tracking
- Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Automated Construction of Artificial Lattice Structures with Designer Electronic States
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- EPANet: Efficient Path Aggregation Network for Underwater Fish Detection
- LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
- YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection
- Recognizing Actions from Robotic View for Natural Human-Robot Interaction
Related