YOLOv12: Attention-Centric Real-Time Object Detectors
2025/02/18 by Tian, Yunjie, Ye, Qixiang, Doermann, David · 95 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2502.12524
Abstract
Enhancing the network architecture of the YOLO framework has been crucial for a long time, but has focused on CNN-based improvements despite the proven superiority of attention mechanisms in modeling capabilities. This is because attention-based models cannot match the speed of CNN-based models. This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. YOLOv12 surpasses all popular real-time object detectors in accuracy with competitive speed. For example, YOLOv12-N achieves 40.6% mAP with an inference latency of 1.64 ms on a T4 GPU, outperforming advanced YOLOv10-N / YOLOv11-N by 2.1%/1.2% mAP with a comparable speed. This advantage extends to other model scales. YOLOv12 also surpasses end-to-end real-time detectors that improve DETR, such as RT-DETR / RT-DETRv2: YOLOv12-S beats RT-DETR-R18 / RT-DETRv2-R18 while running 42% faster, using only 36% of the computation and 45% of the parameters. More comparisons are shown in Figure 1.
Cited by
- An AI-Powered Autonomous Underwater System for Sea Exploration and Scientific Research
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Tiny-YOLOSAM: Fast Hybrid Image Segmentation
- TracktorLive: an integrated real-time object tracking and response system
- Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
- PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
- Generating Risky Samples with Conformity Constraints via Diffusion Models
- YolovN-CBi: A Lightweight and Efficient Architecture for Real-Time Detection of Small UAVs
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- Uni-Parser Technical Report
- Isolated Sign Language Recognition with Segmentation and Pose Estimation
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- TCLeaf-Net: a transformer-convolution framework with global-local attention for robust in-field lesion-level plant leaf disease detection
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial Images
- Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
- Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- Design and Evaluation of a Multi-Agent Perception System for Autonomous Flying Networks
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- Intelligent Image Search Algorithms Fusing Visual Large Models
- A Tri-Modal Dataset and a Baseline System for Tracking Unmanned Aerial Vehicles
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
- SAE-MCVT: A Real-Time and Scalable Multi-Camera Vehicle Tracking Framework Powered by Edge Computing
- Facial Expression Recognition with YOLOv11 and YOLOv12: A Comparative Study
- Hand Held Multi-Object Tracking Dataset in American Football
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- FrameShield: Adversarially Robust Video Anomaly Detection
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- EdgeNavMamba: Mamba Optimized Object Detection for Energy Efficient Edge Devices
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- Counting Hallucinations in Diffusion Models
- MRS-YOLO Railroad Transmission Line Foreign Object Detection Based on Improved YOLO11 and Channel Pruning
- SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
- LadderMoE: Ladder-Side Mixture of Experts Adapters for Bronze Inscription Recognition
- MaizeStandCounting (MaSC): Automated and Accurate Maize Stand Counting from UAV Imagery Using Image Processing and Deep Learning
- OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- In-Field Mapping of Grape Yield and Quality with Illumination-Invariant Deep Learning
- From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
- Beyond Overall Accuracy: Pose- and Occlusion-driven Fairness Analysis in Pedestrian Detection for Autonomous Driving
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- TY-RIST: Tactical YOLO Tricks for Real-time Infrared Small Target Detection
- HyCoVAD: A Hybrid SSL-LLM Model for Complex Video Anomaly Detection
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Real-Time Object Detection Meets DINOv3
- Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection
- A Comparative Benchmark of Real-time Detectors for Blueberry Detection towards Precision Orchard Management
- A Comprehensive Evaluation of YOLO-based Deer Detection Performance on Edge Devices
- Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- RiO-DETR: DETR for Real-time Oriented Object Detection
- YOLO-LAN: Precise Polyp Detection via Optimized Loss, Augmentations and Negatives
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- M3-GloDets: Multi-Region and Multi-Scale Analysis of Fine-Grained Diseased Glomerular Detection
- Degradation-Aware All-in-One Image Restoration via Latent Prior Encoding
- An Empirical Study on the Robustness of YOLO Models for Underwater Object Detection
- MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
- AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
- Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN
- GLIDE: A Coordinated Aerial-Ground Framework for Search and Rescue in Unknown Environments
- A Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset
- Cott-ADNet: Lightweight Real-Time Cotton Boll and Flower Detection Under Field Conditions
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Robust Pan-Cancer Mitotic Figure Detection with YOLOv12
- BuzzSet v1.0: A Dataset for Pollinator Detection in Field Conditions
- OpenTie: Open-vocabulary Sequential Rebar Tying System
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- A Real-Time Diminished Reality Approach to Privacy in MR Collaboration
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring
- Inter-Class Relational Loss for Small Object Detection: A Case Study on License Plates
- A Novel Attention-Augmented Wavelet YOLO System for Real-time Brain Vessel Segmentation on Transcranial Color-coded Doppler
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- DexFruit: Dexterous Manipulation and Gaussian Splatting Inspection of Fruit
- Benchmarking Deep Learning-Based Object Detection Models on Feature Deficient Astrophotography Imagery Dataset
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
- The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
- YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection
- A Multi-Scale Attention-Enhanced Architecture for Gravity Wave Localization in Satellite Imagery
Related