You Only Look Once: Unified, Real-Time Object Detection
2015/06/08 by Joseph Redmon, Redmon, Joseph, Santosh Divvala +5 · 4 voices · 460 citations
#cs.CV
paper · pdf · doi:10.48550/arxiv.1506.02640
Abstract
We present YOLO, a new approach to object detection. Prior work on object detection repurposes classifiers to perform detection. Instead, we frame object detection as a regression problem to spatially separated bounding boxes and associated class probabilities. A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation. Since the whole detection pipeline is a single network, it can be optimized end-to-end directly on detection performance. Our unified architecture is extremely fast. Our base YOLO model processes images in real-time at 45 frames per second. A smaller version of the network, Fast YOLO, processes an astounding 155 frames per second while still achieving double the mAP of other real-time detectors. Compared to state-of-the-art detection systems, YOLO makes more localization errors but is far less likely to predict false detections where nothing exists. Finally, YOLO learns very general representations of objects. It outperforms all other detection methods, including DPM and R-CNN, by a wide margin when generalizing from natural images to artwork on both the Picasso Dataset and the People-Art Dataset.
Cited by
- RISE: Single Static Radar-based Indoor Scene Understanding
- Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models
- Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n
- HGeo-TopoMap: Boosting Topological Mapping with Hierarchical Geometric Priors
- Real-Time EEG Cap Electrode Detection for Guided Point-of-Care Placement
- SpikingMOT: A Spike-Driven Multi-Object Tracker
- SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection
- When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models
- TargetFinder: Detecting Widgets from Pixels on Desktop Interfaces
- Open-Vocabulary Gaze Object Prediction: Benchmark and Method
- Sarus: Privacy-Preserving Multi-Vendor Perception Fusion via Homomorphic Encryption
- Attention from Above: A Multimodal Model for Drone-Based Object Localization
- Technical Design Review of Duke Robotics Club's Oogway & Crush: AUVs for RoboSub 2026
- Toward Optimal Adenovirus Detection Using YOLO26
- IoUCert: Robustness Verification for Anchor-based Object Detectors
- UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation
- Cognitive-YOLO: LLM-Driven Architecture Synthesis from First Principles of Data for Object Detection
- Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking
- Embodied Active Learning under Limited Annotation and Navigation Budget for Object Detection
- Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
- Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method
- Cotton-SF YOLO: Learning Structural and Frequency Cues for Early Cotton Square Detection in Complex Field Environments
- LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data
- Towards Hierarchical Structure Understanding of Newspaper Images
- Moving Like a Human: Ego-Motion-Normalized Temporal Signatures for Real-Time Aerial Person Tracking on Milliwatt-Class Hardware
- Micron-Scale Technosignatures: How a Cubic Metre of Lunar Regolith May Begin to Constrain the Number of Past Technological Civilisations in the Galaxy
- Real-time graph neural networks on FPGAs for the Belle II electromagnetic calorimeter
- GrimACE: automated, multimodal cage-side assessment of pain and well-being in mice
- Comparative assessment of automated and manual monitoring in comprehensive plant–pollinator communities
- Video models are zero-shot learners and reasoners
- Disentangling the Factors of Convergence between Brains and Computer Vision Models
- Advancements in automated nuclei segmentation for histopathology using you only look once-driven approaches: A systematic review
- PCR-ORB: Enhanced ORB-SLAM3 with Point Cloud Refinement Using Deep Learning-Based Dynamic Object Filtering
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks
- Vehicle-to-Everything Cooperative Perception for Autonomous Driving
- A transformer-based multi-stream approach for isolated iranian sign language recognition
- Reading Legends on Ancient Coins: An Object Detection Approach for Character Recognition on a Novel Roman Republican Dataset
- GSpyNetTree-O4: an event validation tool used in the fourth LIGO-Virgo-KAGRA observing run
- Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
- Deep Learning Pose Estimation for Multi-Label Recognition of Combined Hyperkinetic Movement Disorders
- Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- ARCANE–Early Detection of Interplanetary Coronal Mass Ejections
- Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- Granular-ball Guided Masking: Structure-aware Data Augmentation
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Real-World Adversarial Attacks on RF-Based Drone Detectors
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- Progressive Learned Image Compression for Machine Perception
- Gaussian Process Assisted Meta-learning for Image Classification and Object Detection Models
- Long-Range depth estimation using learning based Hybrid Distortion Model for CCTV cameras
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Application of deep learning approaches for medieval historical documents transcription
- Building UI/UX Dataset for Dark Pattern Detection and YOLOv12x-based Real-Time Object Recognition Detection System
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection
- Next-Generation License Plate Detection and Recognition System using YOLOv8
- FlowDet: Unifying Object Detection and Generative Transport Flows
- YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images
- Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics
- MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
- Generative Spatiotemporal Data Augmentation
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- Maritime object classification with SAR imagery using quantum kernel methods
- Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
- Reliable Detection of Minute Targets in High-Resolution Aerial Imagery across Temporal Shifts
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Development and Testing for Perception Based Autonomous Landing of a Long-Range QuadPlane
- Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- SCU-CGAN: Enhancing Fire Detection through Synthetic Fire Image Generation and Dataset Augmentation
- Closed-Loop Robotic Manipulation of Transparent Substrates for Self-Driving Laboratories using Deep Learning Micro-Error Correction
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Relational Visual Similarity
- Improving action classification with brain-inspired deep networks
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- A graph generation pipeline for critical infrastructures based on heuristics, images and depth data
- ShadowWolf -- Automatic Labelling, Evaluation and Model Training Optimised for Camera Trap Wildlife Images
- Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection
- An Integrated System for WEEE Sorting Employing X-ray Imaging, AI-based Object Detection and Segmentation, and Delta Robot Manipulation
- A Comprehensive Framework for Automated Quality Control in the Automotive Industry
- A Hyperspectral Imaging Guided Robotic Grasping System
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- OnSight Pathology: A real-time platform-agnostic computational pathology companion for histopathology
- Out-of-the-box: Black-box Causal Attacks on Object Detectors
- FeatureLens: A Highly Generalizable and Interpretable Framework for Detecting Adversarial Examples Based on Image Features
- Data-Centric Visual Development for Self-Driving Labs
- Is Image-based Object Pose Estimation Ready to Support Grasping?
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- Sleep Apnea Detection on a Wireless Multimodal Wearable Device Without Oxygen Flow Using a Mamba-based Deep Learning Approach
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- Hierarchical Feature Integration for Multi-Signal Automatic Modulation Recognition
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- LLM-Empowered Event-Chain Driven Code Generation for ADAS in SDV systems
- SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- AI/ML Model Cards in Edge AI Cyberinfrastructure: towards Agentic AI
- MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- Video Object Recognition in Mobile Edge Networks: Local Tracking or Edge Detection?
- Intelligent Image Search Algorithms Fusing Visual Large Models
- AIRHILT: A Human-in-the-Loop Testbed for Multimodal Conflict Detection in Aviation
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
- DE-KAN: A Kolmogorov Arnold Network with Dual Encoder for accurate 2D Teeth Segmentation
- Stro-VIGRU: Defining the Vision Recurrent-Based Baseline Model for Brain Stroke Classification
- VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
- EndoSight AI: Deep Learning-Driven Real-Time Gastrointestinal Polyp Detection and Segmentation for Enhanced Endoscopic Diagnostics
- A lightweight detector for real-time detection of remote sensing images
- MfNeuPAN: Proactive End-to-End Navigation in Dynamic Environments via Direct Multi-Frame Point Constraints
- GSpyNetTreeS: a machine learning solution for glitch localization in time and frequency
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- Integrating Deep Learning and Spatial Statistics in Marine Ecosystem Monitoring
- StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- Neural Posterior Estimation with Autoregressive Tiling for Detecting Objects in Astronomical Images
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection
- MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
- Self-supervised learning for multi-label sewer defect classification
- EREBUS: End-to-end Robust Event Based Underwater Simulation
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Enhancing Road Safety Through Multi-Camera Image Segmentation with Post-Encroachment Time Analysis
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- AutoPollS: A tool for automated monitoring of pollinators using deep learning
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
- HideAndSeg: an AI-based tool with automated prompting for octopus segmentation in natural habitats
- Unsupervised Learning for Industrial Defect Detection: A Case Study on Shearographic Data
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- MOBIUS: A Multi-Modal Bipedal Robot that can Walk, Crawl, Climb, and Roll
- A Hybrid YOLOv5-SSD IoT-Based Animal Detection System for Durian Plantation Protection
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- Sewer pipeline condition assessment and defect detection using computer vision
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- MapSAM2: Adapting SAM2 for Automatic Segmentation of Historical Map Images and Time Series
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- Confined Space Underwater Positioning Using Collaborative Robots
- Information-Theoretic Greedy Layer-wise Training for Traffic Sign Recognition
- Jet drop production from bubbles with neighbors
- NOMAD -- Navigating Optimal Model Application to Datastreams
- Gaussian Combined Distance: A Generic Metric for Object Detection
- Improving Classification of Occluded Objects through Scene Context
- Heuristic Adaptation of Potentially Misspecified Domain Support for Likelihood-Free Inference in Stochastic Dynamical Systems
- Sketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Prediction
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
- MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
- Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Impact of purple nutsedge ( Cyperus rotundus ) density on detection accuracy of a YOLOv8 model
- Quantification of plant trait data from herbarium scans in the DiSSCo Research Infrastructure
- iWatchRoadv2: Pothole Detection, Geospatial Mapping, and Intelligent Road Governance
- Mask-Robust Face Verification for Online Learning via YOLOv5 and Residual Networks
- Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- KongNet: A Multi-headed Deep Learning Model for Detection and Classification of Nuclei in Histopathology Images
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- Doubly Smoothed Density Estimation with Application on Miners' Unsafe Act Detection
- GRAP-MOT: Unsupervised Graph-based Position Weighted Person Multi-camera Multi-object Tracking in a Highly Congested Space
- AI Powered Urban Green Infrastructure Assessment Through Aerial Imagery of an Industrial Township
- Synthetic Data for Robust Runway Detection
- Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots
- Automated Morphological Analysis of Neurons in Fluorescence Microscopy Using YOLOv8
- Multi-Camera Worker Tracking in Logistics Warehouse Considering Wide-Angle Distortion
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- Online Object-Level Semantic Mapping for Quadrupeds in Real-World Environments
- Big Data, Tiny Targets: An Exploratory Study in Machine Learning-enhanced Detection of Microplastic from Filters
- Machine Vision-Based Surgical Lighting System:Design and Implementation
- Investigating Adversarial Robustness against Preprocessing used in Blackbox Face Recognition
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- EdgeNavMamba: Mamba Optimized Object Detection for Energy Efficient Edge Devices
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- A Modular Object Detection System for Humanoid Robots Using YOLO
- Detect Anything via Next Point Prediction
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images
- Source-Free Object Detection with Detection Transformer
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Uncertainty-Aware Post-Detection Framework for Enhanced Fire and Smoke Detection in Compact Deep Learning Models
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
- TARO: Toward Semantically Rich Open-World Object Detection
- Mozart: A Chiplet Ecosystem-Accelerator Codesign Framework for Composable Bespoke Application Specific Integrated Circuits
- How Scale Breaks "Normalized Stress" and KL Divergence: Rethinking Quality Metrics
- Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
- Image-based recognition using advanced neural networks can aid surveillance of Agrilus jewel beetles
- LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
- GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Anomaly-Aware YOLO: A Frugal yet Robust Approach to Infrared Small Target Detection
- Comparative Analysis of YOLOv5, Faster R-CNN, SSD, and RetinaNet for Motorbike Detection in Kigali Autonomous Driving Context
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- Road Damage and Manhole Detection using Deep Learning for Smart Cities: A Polygonal Annotation Approach
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- A Hierarchical Agentic Framework for Autonomous Drone-Based Visual Inspection
- Stealing AI Model Weights Through Covert Communication Channels
- Beyond Overall Accuracy: Pose- and Occlusion-driven Fairness Analysis in Pedestrian Detection for Autonomous Driving
- A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
- Using Images from a Video Game to Improve the Detection of Truck Axles
- YOLO-Based Defect Detection for Metal Sheets
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- "Where Can I Park?" Understanding Human Perspectives and Scalably Detecting Disability Parking from Aerial Imagery
- A Cartography of Open Collaboration in Open Source AI: Mapping Practices, Motivations, and Governance in 14 Open Large Language Model Projects
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- Comprehensive Benchmarking of YOLOv11 Architectures for Scalable and Granular Peripheral Blood Cell Detection
- Curriculum Imitation Learning of Distributed Multi-Robot Policies
- UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
- CrashSplat: 2D to 3D Vehicle Damage Segmentation in Gaussian Splatting
- TRAX: TRacking Axles for Accurate Axle Count Estimation
- Advances and Challenges in Machine Learning‐based Image Analysis for Monitoring and Predicting Organic Crystal Formation
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Stochastic activations
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Task-Oriented Computation Offloading for Edge Inference: An Integrated Bayesian Optimization and Deep Reinforcement Learning Framework
- Real-Time Object Detection Meets DINOv3
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
- A Comprehensive Evaluation of YOLO-based Deer Detection Performance on Edge Devices
- You Only Measure Once: On Designing Single-Shot Quantum Machine Learning Models
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- High Clockrate Free-space Optical In-Memory Computing
- Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
- Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
- RiO-DETR: DETR for Real-time Oriented Object Detection
- Enhanced Drift-Aware Computer Vision Architecture for Autonomous Driving
- A holistic perception system of internal and external monitoring for ground autonomous vehicles: AutoTRUST paradigm
- Optimised light sources reduce nocturnal insect attraction at railway stations
- Biological characterization of a mid‐water salinity maximum intrusion over the Northeast US Shelf
- HitoMi-Cam: A Shape-Agnostic Person Detection Method Using the Spectral Characteristics of Clothing
- From Unstable to Playable: Stabilizing Angry Birds Levels via Object Segmentation
- YOLO-LAN: Precise Polyp Detection via Optimized Loss, Augmentations and Negatives
- Spectral Signature Mapping from RGB Imagery for Terrain-Aware Navigation
- Investigating Traffic Accident Detection Using Multimodal Large Language Models
- MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
- Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training
- Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- Few-Shot Pattern Detection via Template Matching and Regression
- Automated Coral Spawn Monitoring for Reef Restoration: The Coral Spawn and Larvae Imaging Camera System (CSLICS)
- MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
- Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models
- Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
- GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
- Incorporating the Refractory Period into Spiking Neural Networks through Spike-Triggered Threshold Dynamics
- LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object Detection
- Enhanced Detection of Tiny Objects in Aerial Images
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- Saccadic Vision for Fine-Grained Visual Classification
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Synthetic-to-Real Object Detection using YOLOv11 and Domain Randomization Strategies
- A Synthetic Dataset for Manometry Recognition in Robotic Applications
- Efficient 3D Perception on Embedded Systems via Interpolation-Free Tri-Plane Lifting and Volume Fusion
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- GW-YOLO: Multi-transient segmentation in LIGO using computer vision
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Deep Learning for Analyzing Chaotic Dynamics in Biological Time Series: Insights from Frog Heart Signals
- Federated Reinforcement Learning for Runtime Optimization of AI Applications in Smart Eyewears
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- A Synthetic Data Pipeline for Supporting Manufacturing SMEs in Visual Assembly Control
- From Pixels to Shelf: End-to-End Algorithmic Control of a Mobile Manipulator for Supermarket Stocking and Fronting
- CSIYOLO: An Intelligent CSI-based Scatter Sensing Framework for Integrated Sensing and Communication Systems
- Advanced Layout Analysis Models for Docling
- Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
- Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
- Scensory: Automated Real-Time Fungal Identification and Spatial Mapping
- A Co-Training Semi-Supervised Framework Using Faster R-CNN and YOLO Networks for Object Detection in Densely Packed Retail Images
- Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection
- IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- Studying collective animal behaviour with drones and computer vision
- MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis
- BatStation: Toward In-Situ Radar Sensing on 5G Base Stations with Zero-Shot Template Generation
- Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
- Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
- Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- Multi-Modal Camera-Based Detection of Vulnerable Road Users
- Enhancing Low-Altitude Airspace Security: MLLM-Enabled UAV Intent Recognition
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- A New Hybrid Model of Generative Adversarial Network and You Only Look Once Algorithm for Automatic License-Plate Recognition
- Exploring Light-Weight Object Recognition for Real-Time Document Detection
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Towards Open World Detection: A Survey
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- treeX: Unsupervised Tree Instance Segmentation in Dense Forest Point Clouds
- Self-Organizing Aerial Swarm Robotics for Resilient Load Transportation : A Table-Mechanics-Inspired Approach
- VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- An Investigation of Visual Foundation Models Robustness
- Self-Validated Learning for Particle Separation: A Correctness-Based Self-Training Framework Without Human Labels
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
- High-Precision Mixed Feature Fusion Network Using Hypergraph Computation for Cervical Abnormal Cell Detection
- RT-DETRv2 Explained in 8 Illustrations
- A Real-Time, Vision-Based System for Badminton Smash Speed Estimation on Mobile Devices
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Artificial intelligence and companion animals: Perspectives on digital healthcare for dogs, cats, and pet ownership
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- End-to-End Analysis of Charge Stability Diagrams with Transformers
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- Quantization Robustness to Input Degradations for Object Detection
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Mini-Batch Robustness Verification of Deep Neural Networks
- Robust and Label-Efficient Deep Waste Detection
- Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- BirdRecorder's AI on Sky: Safeguarding birds of prey by detection and classification of tiny objects around wind turbines
- AQ-PCDSys: An Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration
- A Real-Time Diminished Reality Approach to Privacy in MR Collaboration
- Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring
- Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels
- SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Safe and Transparent Robots for Human-in-the-Loop Meat Processing
- Improved Mapping Between Illuminations and Sensors for RAW Images
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Online 3D Gaussian Splatting Modeling with Novel View Selection
- Fracture Detection and Localisation in Wrist and Hand Radiographs using Detection Transformer Variants
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- Data Shift of Object Detection in Autonomous Driving
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- An Exploratory Study on Crack Detection in Concrete through Human-Robot Collaboration
- Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
- Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
- Visuomotor Grasping with World Models for Surgical Robots
- Colon Polyps Detection from Colonoscopy Images Using Deep Learning
- Lightweight CNNs for Embedded SAR Ship Target Detection and Classification
- CSNR and JMIM Based Spectral Band Selection for Reducing Metamerism in Urban Driving
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
- IPG: Incremental Patch Generation for Generalized Adversarial Patch Training
- MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities
- Real Time Child Abduction And Detection System
- StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
- NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3D Gaussian Reconstruction
- DexFruit: Dexterous Manipulation and Gaussian Splatting Inspection of Fruit
- Automated detection of fin whales with distributed acoustic sensing in the Arctic and Mediterranean
- CountQA: How Well Do MLLMs Count in the Wild?
- Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a User-Centric Optimization Space
- Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
- SGDFuse: SAM-Guided Diffusion for High-Fidelity Infrared and Visible Image Fusion
- RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
- ULU: A Unified Activation Function
- CSRAP: Enhanced Canvas Attention Scheduling for Real-Time Mission Critical Perception
- Drone Detection with Event Cameras
- Learning Using Privileged Information for Litter Detection
- CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
- Deep learning framework for crater detection and identification on the Moon and Mars
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- An RGB-D Camera-Based Multi-Small Flying Anchors Control for Wire-Driven Robots Connecting to the Environment
- Understanding the Risks of Asphalt Art on the Reliability of Surveillance Perception Systems
- MetAdv: A Unified and Interactive Adversarial Testing Platform for Autonomous Driving
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Spatial-Frequency Aware for Object Detection in RAW Image
- Human-Robot Red Teaming for Safety-Aware Reasoning
- Fusion Sampling Validation in Data Partitioning for Machine Learning
- AI-Driven Collaborative Satellite Object Detection for Space Sustainability
- Backdoor Attacks on Deep Learning Face Detection
- Towards Field-Ready AI-based Malaria Diagnosis: A Continual Learning Approach
- Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
- YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection
- SpectraSentinel: LightWeight Dual-Stream Real-Time Drone Detection, Tracking and Payload Identification
- Object Recognition Datasets and Challenges: A Review
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Knowledge Augmentation via Synthetic Data: A Framework for Real-World ECG Image Classification
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors
- SARD: A YOLOv8-Based System for Solar Active Region Detection with SDO/HMI Magnetograms
- Low-Cost Test-Time Adaptation for Robust Video Editing
- Tracking Moose using Aerial Object Detection
- Customize Multi-modal RAI Guardrails with Precedent-based predictions
- Region-based Cluster Discrimination for Visual Representation Learning
- DriveIndia: An Object Detection Dataset for Diverse Indian Traffic Scenes
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- Efficient Self-Supervised Neuro-Analytic Visual Servoing for Real-time Quadrotor Control
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- JDATT: A Joint Distillation Framework for Atmospheric Turbulence Mitigation and Target Detection
- Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and beyond
- Underwater Waste Detection Using Deep Learning A Performance Comparison of YOLOv7 to 10 and Faster RCNN
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- CDA-SimBoost: A Unified Framework Bridging Real Data and Simulation for Infrastructure-Based CDA Systems
- TCM-Tongue: A Standardized Tongue Image Dataset with Pathological Annotations for AI-Assisted TCM Diagnosis
- Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
- Real-Time Object Detection and Classification using YOLO for Edge FPGAs
- Towards Large Scale Geostatistical Methane Monitoring with Part-based Object Detection
- Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
- You Only Look Once [wikipedia]
- Small object detection [wikipedia]
- Image analysis [wikipedia]
- Object detection [wikipedia]
Discussions
- You Only Look Once: Unified, Real-Time Object Detection [hn, 64 points, 8 comments]
- I think people mostly use a YOLO architecture for this. Original paper: arxiv.org/abs/1506.02640) [bsky, 3 points, 1 comments]
- You Only Look Once: Unified, Real-Time Object Detection [hn, 1 points, 0 comments]
- Because YOLO (You Only Look Once)! Is the first time I hear about this algorithm <3! arxiv.org/pdf/1506.026... [bsky, 0 points, 0 comments]
Related