YOLOv3: An Incremental Improvement
2018/04/08 by Joseph Redmon, Ali Farhadi, Redmon, Joseph +1 · 9 voices · 389 citations
Computer Science · Medicine · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Retinal Imaging and Analysis
paper · pdf · doi:10.48550/arxiv.1804.02767
Abstract
We present some updates to YOLO! We made a bunch of little design changes to make it better. We also trained this new network that's pretty swell. It's a little bigger than last time but more accurate. It's still fast though, don't worry. At 320x320 YOLOv3 runs in 22 ms at 28.2 mAP, as accurate as SSD but three times faster. When we look at the old .5 IOU mAP detection metric YOLOv3 is quite good. It achieves 57.9 mAP@50 in 51 ms on a Titan X, compared to 57.5 mAP@50 in 198 ms by RetinaNet, similar performance but 3.8x faster. As always, all the code is online at https://pjreddie.com/yolo/
Cited by
- SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection
- A real-time RGB-D perception pipeline for autonomous impact hammers in mining: self-filtering, rock segmentation and rock-breaking poses generation
- Open-Vocabulary Gaze Object Prediction: Benchmark and Method
- Attention from Above: A Multimodal Model for Drone-Based Object Localization
- IoUCert: Robustness Verification for Anchor-based Object Detectors
- AdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature Suppression
- Thermal Topology Collapse: Universal Physical Patch Attacks on Infrared Vision Systems
- Embodied Active Learning under Limited Annotation and Navigation Budget for Object Detection
- Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- Recursive Language Models
- FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
- High-resolution efficient image generation from WiFi CSI using a pretrained latent diffusion model
- Advancements in automated nuclei segmentation for histopathology using you only look once-driven approaches: A systematic review
- BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Evaluating an Adaptive Multispectral Turret System for Autonomous Tracking Across Variable Illumination Conditions
- PD-SORT: Occlusion-Robust Multi-Object Tracking Using Pseudo-Depth Cues
- Reading Legends on Ancient Coins: An Object Detection Approach for Character Recognition on a Novel Roman Republican Dataset
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images
- A Comprehensive Safety Metric to Evaluate Perception in Autonomous Systems
- TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
- CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- TCLeaf-Net: a transformer-convolution framework with global-local attention for robust in-field lesion-level plant leaf disease detection
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
- Efficient Feature Compression for Machines with Global Statistics Preservation
- Visual Heading Prediction for Autonomous Aerial Vehicles
- Attention-based Joint Detection of Object and Semantic Part
- A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
- New VVC profiles targeting Feature Coding for Machines
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- Density Map Guided Object Detection in Aerial Images
- A Comparative Review of Recent Few-Shot Object Detection Algorithms
- DPNET: Dual-Path Network for Efficient Object Detectioj with Lightweight Self-Attention
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Object Detection Based Handwriting Localization
- Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
- Turkey Behavior Identification System with a GUI Using Deep Learning and Video Analytics
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- Uncertainty-Aware Model Adaptation for Unsupervised Cross-Domain Object Detection
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- RPCL: A Framework for Improving Cross-Domain Detection with Auxiliary Tasks
- Data-Driven Multi-Emitter Localization Using Spatially Distributed Power Measurements
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
- Small Object Detection for Birds with Swin Transformer
- Power-Efficient Autonomous Mobile Robots
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Deep Learning in Computer-Aided Diagnosis and Treatment of Tumors: A Survey
- A Unified Framework of Constrained Robust Submodular Optimization with\n Applications
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- Machine Learning Techniques to Detect and Characterise Whistler Radio\n Waves
- GSpyNetTreeS: a machine learning solution for glitch localization in time and frequency
- Unbiased Teacher for Semi-Supervised Object Detection
- T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Gliding Vertex on the Horizontal Bounding Box for Multi-Oriented Object Detection
- Physically Realistic Sequence-Level Adversarial Clothing for Robust Human-Detection Evasion
- Orthographic Feature Transform for Monocular 3D Object Detection
- PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation
- COMET: Context-Aware IoU-Guided Network for Small Object Tracking
- Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
- XNOR-Net++: Improved Binary Neural Networks
- Multi-Scale Progressive Fusion Network for Single Image Deraining
- Serial-parallel Multi-Scale Feature Fusion for Anatomy-Oriented Hand Joint Detection
- Towards Real-Time Multi-Object Tracking
- LiveMap: Real-Time Dynamic Map in Automotive Edge Computing
- Inter-Homines: Distance-Based Risk Estimation for Human Safety
- Objects detection for remote sensing images based on polar coordinates
- Comparison of object detection methods for crop damage assessment using deep learning
- Privacy Aware Person Detection in Surveillance Data
- Measurement-driven Analysis of an Edge-Assisted Object Recognition System
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
- Learning Fourier shapes to probe the geometric world of deep neural networks
- Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features
- Gaussian Combined Distance: A Generic Metric for Object Detection
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Dual Refinement Feature Pyramid Networks for Object Detection
- CityFlow-NL: Tracking and Retrieval of Vehicles at City Scale by Natural Language Descriptions
- Confluence: A Robust Non-IoU Alternative to Non-Maxima Suppression in Object Detection
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
- On Physical Adversarial Patches for Object Detection
- Red Blood Cell Segmentation with Overlapping Cell Separation and Classification on Imbalanced Dataset
- Explicit Shape Encoding for Real-Time Instance Segmentation
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- D-VPnet: A Network for Real-time Dominant Vanishing Point Detection in Natural Scenes
- Chargrid-OCR: End-to-end Trainable Optical Character Recognition for Printed Documents using Instance Segmentation
- Computer Stereo Vision for Autonomous Driving
- DEFT: Detection Embeddings for Tracking
- Object Detection in 20 Years: A Survey
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- Sequence-SOD: Bio-inspired Sequence-aware Spiking Object Detection for Event Cameras
- Angel's Girl for Blind Painters: an Efficient Painting Navigation System Validated by Multimodal Evaluation Approach
- Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction
- Query by Semantic Sketch
- Comparison of Object Detection Algorithms Using Video and Thermal Images Collected from a UAS Platform: An Application of Drones in Traffic Management
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- A Dataset and Application for Facial Recognition of Individual Gorillas in Zoo Environments
- JRDB: A Dataset and Benchmark of Egocentric Robot Visual Perception of Humans in Built Environments
- FOD-A: A Dataset for Foreign Object Debris in Airports
- Pixel-Semantic Revise of Position Learning A One-Stage Object Detector\n with A Shared Encoder-Decoder
- Rethinking Classification and Localization for Object Detection
- Weakly supervised one-stage vision and language disease detection using large scale pneumonia and pneumothorax studies
- Monocular, One-stage, Regression of Multiple 3D People
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Consistent Optimization for Single-Shot Object Detection
- PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
- Testing Deep Learning Models for Image Analysis Using Object-Relevant Metamorphic Relations
- A Fusion Adversarial Underwater Image Enhancement Network with a Public Test Dataset
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Over-sampling De-occlusion Attention Network for Prohibited Items Detection in Noisy X-ray Images
- Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation
- Multiple Object Tracking with Correlation Learning
- End-to-end Active Object Tracking and Its Real-world Deployment via Reinforcement Learning
- From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models
- Substitute Teacher Networks: Learning with Almost No Supervision
- Bag of Freebies for Training Object Detection Neural Networks
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- SurveilEdge: Real-time Video Query based on Collaborative Cloud-Edge Deep Learning
- Dynamic Region Division for Adaptive Learning Pedestrian Counting
- VarifocalNet: An IoU-aware Dense Object Detector
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Detection as Regression: Certified Object Detection by Median Smoothing
- EfficientDet: Scalable and Efficient Object Detection
- A Survey of Fish Tracking Techniques Based on Computer Vision
- Dynamic Adversarial Patch for Evading Object Detection Models
- Learning a Proposal Classifier for Multiple Object Tracking
- LLVIP: A Visible-infrared Paired Dataset for Low-light Vision
- Toward Intelligent Sensing: Intermediate Deep Feature Compression
- Pareto-Optimal Bit Allocation for Collaborative Intelligence
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- Vote from the Center: 6 DoF Pose Estimation in RGB-D Images by Radial Keypoint Voting
- Decision-based Universal Adversarial Attack
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Real-Time target detection in maritime scenarios based on YOLOv3 model
- Towards More Flexible and Accurate Object Tracking with Natural Language: Algorithms and Benchmark
- Non-local RoI for Cross-Object Perception
- Occluded Prohibited Items Detection: an X-ray Security Inspection Benchmark and De-occlusion Attention Module
- A Real-time Global Inference Network for One-stage Referring Expression Comprehension
- Multimodal Generative Models for Compositional Representation Learning
- A review of uncertainty quantification in deep learning: Techniques, applications and challenges
- Distilling Image Classifiers in Object Detectors
- FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
- Resource-Constrained Simultaneous Detection and Labeling of Objects in High-Resolution Satellite Images
- Saber, opinion y ciencia, de Daniel Quesada
- A Survey on Sensor Technologies for Unmanned Ground Vehicles
- EDSL: An Encoder-Decoder Architecture with Symbol-Level Features for Printed Mathematical Expression Recognition
- IoU-aware Single-stage Object Detector for Accurate Localization
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis Network
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Detection and Tracking Meet Drones Challenge
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- FVNet: 3D Front-View Proposal Generation for Real-Time Object Detection from Point Clouds
- WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
- General Instance Distillation for Object Detection
- A Practical Framework for ROI Detection in Medical Images -- a case study for hip detection in anteroposterior pelvic radiographs
- ARC: A Vision-based Automatic Retail Checkout System
- AI Pose Analysis and Kinematic Profiling of Range-of-Motion Variations in Resistance Training
- Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts
- AGE Challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
- Confidence Calibration with Bounded Error Using Transformations
- Monitoring Horses in Stalls: From Object to Event Detection
- Deep Knowledge Tracing with Learning Curves
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- A Modular Object Detection System for Humanoid Robots Using YOLO
- NarrationBot and InfoBot: A Hybrid System for Automated Video Description
- Detect Anything via Next Point Prediction
- Collaborative Multi-Robot Systems for Search and Rescue: Coordination and Perception
- A Modular AIoT Framework for Low-Latency Real-Time Robotic Teleoperation in Smart Cities
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Automated Defect Detection for Mass-Produced Electronic Components Based on YOLO Object Detection Models
- A Cost Effective Solution for Road Crack Inspection using Cameras and Deep Neural Networks
- Detection of Dataset Shifts in Learning-Enabled Cyber-Physical Systems using Variational Autoencoder for Regression
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- ThunderNet: Towards Real-time Generic Object Detection
- QANet -- Quality Assurance Network for Image Segmentation
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- Whole Body Model Predictive Control for Spin-Aware Quadrupedal Table Tennis
- Towards Adversarially Robust Object Detection
- Meta Adversarial Training against Universal Patches
- Video Surveillance for Road Traffic Monitoring
- Robust 2D/3D Vehicle Parsing in CVIS
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- EOLO: Embedded Object Segmentation only Look Once
- MaizeStandCounting (MaSC): Automated and Accurate Maize Stand Counting from UAV Imagery Using Image Processing and Deep Learning
- Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
- GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Investigating mixed traffic dynamics of pedestrians and non-motorized vehicles at urban intersections: Observation experiments and modelling
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Skin Lesion Classification Based on ResNet-50 Enhanced With Adaptive Spatial Feature Fusion
- Oriented RepPoints for Aerial Object Detection
- Humanly Certifying Superhuman Classifiers
- Using Images from a Video Game to Improve the Detection of Truck Axles
- YOLO-Based Defect Detection for Metal Sheets
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- Towards Palmprint Verification On Smartphones
- EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments
- Voxel-FPN: multi-scale voxel feature aggregation in 3D object detection from point clouds
- Machine Learning Based Analysis of Finnish World War II Photographers
- DetectorGuard: Provably Securing Object Detectors against Localized Patch Hiding Attacks
- Visual Relationship Detection using Scene Graphs: A Survey
- Scope Head for Accurate Localization in Object Detection
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Deep-RBF Networks Revisited: Robust Classification with Rejection
- Resolving Class Imbalance in Object Detection with Weighted Cross Entropy Losses
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
- MM-FSOD: Meta and metric integrated few-shot object detection
- NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination
- (AF)2-S3Net: Attentive Feature Fusion with Adaptive Feature Selection for Sparse Semantic Segmentation Network
- Dynamic DNN Decomposition for Lossless Synergistic Inference
- Camera Exposure Control for Robust Robot Vision with Noise-Aware Image Quality Assessment
- Representation Sharing for Fast Object Detector Search and Beyond
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- SOLO: A Simple Framework for Instance Segmentation
- A Survey of Autonomous Driving: <i>Common Practices and Emerging Technologies</i>
- IMMVP: An Efficient Daytime and Nighttime On-Road Object Detector
- Movement Tracks for the Automatic Detection of Fish Behavior in Videos
- An improved helmet detection method for YOLOv3 on an unbalanced dataset
- Learning Universal Shape Dictionary for Realtime Instance Segmentation
- Video Anomaly Detection by Estimating Likelihood of Representations
- A Comprehensive Review on Recent Methods and Challenges of Video\n Description
- Skeleon-Based Typing Style Learning For Person Identification
- Expedited Multi-Target Search with Guaranteed Performance via Multi-fidelity Gaussian Processes
- Assessing the Alignment of Popular CNNs to the Brain for Valence Appraisal
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- Automated Detecting and Placing Road Objects from Street-level Images
- Trustworthy Convolutional Neural Networks: A Gradient Penalized-based Approach
- Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high resolution RGB labelled images to develop and benchmark wheat head detection methods
- Enhanced Detection of Tiny Objects in Aerial Images
- Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- A MAC-less Neural Inference Processor Supporting Compressed, Variable\n Precision Weights
- Towards an efficient framework for Data Extraction from Chart Images
- Modification method for single-stage object detectors that allows to\n exploit the temporal behaviour of a scene to improve detection accuracy
- Video based real-time positional tracker
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Performance Optimization of YOLO-FEDER FusionNet for Robust Drone Detection in Visually Complex Environments
- A Survey of Deep Learning for Scientific Discovery
- A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
- DeepSperm: A robust and real-time bull sperm-cell detection in densely populated semen videos
- An Orientation Factor for Object-Oriented SLAM
- Trained Model Fusion for Object Detection using Gating Network
- Unifying Variational Inference and PAC-Bayes for Supervised Learning that Scales
- Automatically detecting pig position and posture by 2D camera imaging and deep learning
- CSIYOLO: An Intelligent CSI-based Scatter Sensing Framework for Integrated Sensing and Communication Systems
- Synthetic Occlusion Augmentation with Volumetric Heatmaps for the 2018 ECCV PoseTrack Challenge on 3D Human Pose Estimation
- Review: deep learning on 3D point clouds
- Rethinking Channel Dimensions for Efficient Model Design
- WebSight: A Vision-First Architecture for Robust Web Agents
- Locate then Segment: A Strong Pipeline for Referring Image Segmentation
- TKD: Temporal Knowledge Distillation for Active Perception
- Research on Fast Text Recognition Method for Financial Ticket Image
- Deep Learning Based FDD Non-Stationary Massive MIMO Downlink Channel Reconstruction
- VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Two-Stage Swarm Intelligence Ensemble Deep Transfer Learning (SI-EDTL) for Vehicle Detection Using Unmanned Aerial Vehicles
- DEPFusion: Dual-Domain Enhancement and Priority-Guided Mamba Fusion for UAV Multispectral Object Detection
- Parse Graph-Based Visual-Language Interaction for Human Pose Estimation
- FCOS: Fully Convolutional One-Stage Object Detection
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Deep learning visual analysis in laparoscopic surgery: a systematic review and diagnostic test accuracy meta-analysis
- Semantic Segmentation for Compound figures
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Detection of E-scooter Riders in Naturalistic Scenes
- FCOS: A simple and strong anchor-free object detector
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Plant diseases and pests detection based on deep learning: a review
- Video Analytics with Zero-streaming Cameras
- Assembly of randomly placed parts realized by using only one robot arm with a general parallel-jaw gripper
- Object Detection in Specific Traffic Scenes using YOLOv2
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
- Efficient Pipelines for Vision-Based Context Sensing
- Global Context Aware RCNN for Object Detection
- Non-local RoIs for Instance Segmentation
- An End-to-End Framework for Video Multi-Person Pose Estimation
- Empowering cyberphysical systems of systems with intelligence
- COVID-19 personal protective equipment detection using real-time deep learning methods
- Probabilistic two-stage detection
- QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in\n Object-oriented SLAM
- YOLOX: Exceeding YOLO Series in 2021
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Objects in Semantic Topology
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Referring Expression Comprehension: A Survey of Methods and Datasets
- Streamlining the Development of Active Learning Methods in Real-World Object Detection
- Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Weed Detection in Challenging Field Conditions: A Semi-Supervised Framework for Overcoming Shadow Bias and Data Scarcity
- Learning and Reasoning with the Graph Structure Representation in Robotic Surgery
- Exploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
- Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Deep Active Learning for Remote Sensing Object Detection
- Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
- Revisiting Knowledge Distillation for Object Detection
- CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
- FS-Net: Fast Shape-based Network for Category-Level 6D Object Pose Estimation with Decoupled Rotation Mechanism
- INSTA-YOLO: Real-Time Instance Segmentation
- G2L-Net: Global to Local Network for Real-time 6D Pose Estimation with Embedding Vector Features
- A Cost-Effective Framework for Predicting Parking Availability Using Geospatial Data and Machine Learning
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- Technical Report: Reactive Semantic Planning in Unexplored Semantic Environments Using Deep Perceptual Feedback
- A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network
- Semi-supervised Image Dehazing via Expectation-Maximization and Bidirectional Brownian Bridge Diffusion Models
- Colon Polyps Detection from Colonoscopy Images Using Deep Learning
- Lameness detection in dairy cows using pose estimation and bidirectional LSTMs
- Towards Powerful and Practical Patch Attacks for 2D Object Detection in Autonomous Driving
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Background Segmentation for Vehicle Re-Identification
- Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- v2e: From Video Frames to Realistic DVS Events
- CPGAN: Full-Spectrum Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Real-time landmark detection for precise endoscopic submucosal\n dissection via shape-aware relation network
- Non-anchor-based vehicle detection for traffic surveillance using bounding ellipses
- One-Shot Instance Segmentation
- Boosting Image Outpainting with Semantic Layout Prediction
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Strawberry Detection Using a Heterogeneous Multi-Processor Platform
- LF-YOLO: A Lighter and Faster YOLO for Weld Defect Detection of X-ray Image
- Domain Adaptive YOLO for One-Stage Cross-Domain Detection
- WDR FACE: The First Database for Studying Face Detection in Wide Dynamic Range
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- LAMP: Large-Scale Autonomous Mapping and Positioning for Exploration of\n Perceptually-Degraded Subterranean Environments
- ULU: A Unified Activation Function
- YOLOv8-Based Deep Learning Model for Automated Poultry Disease Detection and Health Monitoring paper
- Robust Processing-In-Memory Neural Networks via Noise-Aware Normalization
- CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps
- A Smartphone-based System for Real-time Early Childhood Caries Diagnosis
- Deep learning framework for crater detection and identification on the Moon and Mars
- Talking Detection In Collaborative Learning Environments
- Adversarial Attention Perturbations for Large Object Detection Transformers
- The Indirect Convolution Algorithm
- CoFF: Cooperative Spatial Feature Fusion for 3D Object Detection on Autonomous Vehicles
- Few-shot Object Detection with Self-adaptive Attention Network for Remote Sensing Images
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Privacy-Preserving Driver Drowsiness Detection with Spatial Self-Attention and Federated Learning
- "Who is Driving around Me?" Unique Vehicle Instance Classification using Deep Neural Features
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image
- You Only Look Once [wikipedia]
- Small object detection [wikipedia]
Discussions
- Not an essay but I think about this section from the YOLOv3 paper often https://arxiv.org/pdf/1804.02767 [bsky, 74 points, 3 comments]
- how did i miss this paper?? arxiv.org/abs/1804.02767 [bsky, 28 points, 5 comments]
- I’ve read perhaps dozens of CV research papers and found that the most important paragraph, by a wide margin, to be the conclusion of the yoloV3 tech report… arxiv.org/abs/1804.02767 [bsky, 11 points, 0 comments]
- My award for most honest paper goes to: arxiv.org/abs/1804.02767 An arxiv-only paper describing v3 of YOLO, a vision algo that already had two published papers. With huge adoption already, authors wer [bsky, 4 points, 0 comments]
- YOLOv3: An Incremental Improvement [hn, 1 points, 0 comments]
- Today a colleague forwarded this article to me, from Joseph Redmon and Ali Farhadi YOLOv3: An Incremental Improvement arxiv.org/abs/1804.02767 The relevant passage: [bsky, 1 points, 0 comments]
- Some academic writing style palate cleanser for frustrated paper writers: arxiv.org/pdf/1804.02767 [bsky, 1 points, 0 comments]
- Reminded today that this paper exists: arxiv.org/pdf/1804.02767 [bsky, 1 points, 0 comments]
- This guy arxiv.org/abs/1804.02767 [bsky, 1 points, 0 comments]
Related