YOLOv3: An Incremental Improvement
2018/04/08 by Joseph Redmon, Ali Farhadi, Redmon, Joseph +1 · 9 voices · 5,889 citations
Computer Science · Engineering · Medicine · Psychology · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Code (set theory) #Computer science #Computer vision #Engineering #Geology #Metric (unit) #Programming language #Psychology #Retinal Imaging and Analysis #Swell #Titan (rocket family) #Worry #cs.CV
paper · pdf · doi:10.48550/arxiv.1804.02767
published in arXiv (Cornell University) (Cornell University) · Tech Report
arxiv created 2018/04/08 · openalex publication_date 2018/04/08 · arxiv updated 2018/04/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We present some updates to YOLO! We made a bunch of little design changes to make it better. We also trained this new network that's pretty swell. It's a little bigger than last time but more accurate. It's still fast though, don't worry. At 320x320 YOLOv3 runs in 22 ms at 28.2 mAP, as accurate as SSD but three times faster. When we look at the old .5 IOU mAP detection metric YOLOv3 is quite good. It achieves 57.9 mAP@50 in 51 ms on a Titan X, compared to 57.5 mAP@50 in 198 ms by RetinaNet, similar performance but 3.8x faster. As always, all the code is online at https://pjreddie.com/yolo/
Cited by
- SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection
- A real-time RGB-D perception pipeline for autonomous impact hammers in mining: self-filtering, rock segmentation and rock-breaking poses generation
- Open-Vocabulary Gaze Object Prediction: Benchmark and Method
- Attention from Above: A Multimodal Model for Drone-Based Object Localization
- IoUCert: Robustness Verification for Anchor-based Object Detectors
- AdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature Suppression
- Thermal Topology Collapse: Universal Physical Patch Attacks on Infrared Vision Systems
- Embodied Active Learning under Limited Annotation and Navigation Budget for Object Detection
- Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- Recursive Language Models
- FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
- High-resolution efficient image generation from WiFi CSI using a pretrained latent diffusion model
- Advancements in automated nuclei segmentation for histopathology using you only look once-driven approaches: A systematic review
- BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Evaluating an Adaptive Multispectral Turret System for Autonomous Tracking Across Variable Illumination Conditions
- PD-SORT: Occlusion-Robust Multi-Object Tracking Using Pseudo-Depth Cues
- Reading Legends on Ancient Coins: An Object Detection Approach for Character Recognition on a Novel Roman Republican Dataset
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images
- A Comprehensive Safety Metric to Evaluate Perception in Autonomous Systems
- TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
- CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
- LFFD: A Light and Fast Face Detector for Edge Devices
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- TCLeaf-Net: a transformer-convolution framework with global-local attention for robust in-field lesion-level plant leaf disease detection
- LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing
- Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
- Efficient Feature Compression for Machines with Global Statistics Preservation
- Visual Heading Prediction for Autonomous Aerial Vehicles
- Attention-based Joint Detection of Object and Semantic Part
- A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
- New VVC profiles targeting Feature Coding for Machines
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- Density Map Guided Object Detection in Aerial Images
- A Comparative Review of Recent Few-Shot Object Detection Algorithms
- DPNET: Dual-Path Network for Efficient Object Detectioj with Lightweight Self-Attention
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Object Detection Based Handwriting Localization
- Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
- Turkey Behavior Identification System with a GUI Using Deep Learning and Video Analytics
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- Uncertainty-Aware Model Adaptation for Unsupervised Cross-Domain Object Detection
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- RPCL: A Framework for Improving Cross-Domain Detection with Auxiliary Tasks
- Data-Driven Multi-Emitter Localization Using Spatially Distributed Power Measurements
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
- Small Object Detection for Birds with Swin Transformer
- Power-Efficient Autonomous Mobile Robots
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Deep Learning in Computer-Aided Diagnosis and Treatment of Tumors: A Survey
- A Unified Framework of Constrained Robust Submodular Optimization with\n Applications
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- Machine Learning Techniques to Detect and Characterise Whistler Radio\n Waves
- GSpyNetTreeS: a machine learning solution for glitch localization in time and frequency
- Unbiased Teacher for Semi-Supervised Object Detection
- T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Gliding vertex on the horizontal bounding box for multi-oriented object detection
- Physically Realistic Sequence-Level Adversarial Clothing for Robust Human-Detection Evasion
- Orthographic Feature Transform for Monocular 3D Object Detection
- PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation
- COMET: Context-Aware IoU-Guided Network for Small Object Tracking
- Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
- XNOR-Net++: Improved Binary Neural Networks
- Multi-Scale Progressive Fusion Network for Single Image Deraining
- Serial-parallel Multi-Scale Feature Fusion for Anatomy-Oriented Hand Joint Detection
- Towards Real-Time Multi-Object Tracking
- LiveMap: Real-Time Dynamic Map in Automotive Edge Computing
- Inter-Homines: Distance-Based Risk Estimation for Human Safety
- Objects detection for remote sensing images based on polar coordinates
- Comparison of object detection methods for crop damage assessment using deep learning
- Privacy Aware Person Detection in Surveillance Data
- Measurement-driven Analysis of an Edge-Assisted Object Recognition System
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
- Learning Fourier shapes to probe the geometric world of deep neural networks
- Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features
- Gaussian Combined Distance: A Generic Metric for Object Detection
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Dual Refinement Feature Pyramid Networks for Object Detection
- CityFlow-NL: Tracking and Retrieval of Vehicles at City Scale by Natural Language Descriptions
- Confluence: A Robust Non-IoU Alternative to Non-Maxima Suppression in Object Detection
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
- On Physical Adversarial Patches for Object Detection
- Red Blood Cell Segmentation with Overlapping Cell Separation and Classification on Imbalanced Dataset
- Explicit Shape Encoding for Real-Time Instance Segmentation
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- D-VPnet: A Network for Real-time Dominant Vanishing Point Detection in Natural Scenes
- Chargrid-OCR: End-to-end Trainable Optical Character Recognition for Printed Documents using Instance Segmentation
- Computer Stereo Vision for Autonomous Driving
- DEFT: Detection Embeddings for Tracking
- Object Detection in 20 Years: A Survey
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- Sequence-SOD: Bio-inspired Sequence-aware Spiking Object Detection for Event Cameras
- Angel's Girl for Blind Painters: an Efficient Painting Navigation System Validated by Multimodal Evaluation Approach
- Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction
- Query by Semantic Sketch
- Comparison of Object Detection Algorithms Using Video and Thermal Images Collected from a UAS Platform: An Application of Drones in Traffic Management
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- A Dataset and Application for Facial Recognition of Individual Gorillas in Zoo Environments
- JRDB: A Dataset and Benchmark of Egocentric Robot Visual Perception of Humans in Built Environments
- FOD-A: A Dataset for Foreign Object Debris in Airports
- Pixel-Semantic Revise of Position Learning A One-Stage Object Detector with A Shared Encoder-Decoder
- Rethinking Classification and Localization for Object Detection
- Weakly supervised one-stage vision and language disease detection using large scale pneumonia and pneumothorax studies
- Monocular, One-stage, Regression of Multiple 3D People
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Consistent Optimization for Single-Shot Object Detection
- PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
- Testing Deep Learning Models for Image Analysis Using Object-Relevant Metamorphic Relations
- A Fusion Adversarial Underwater Image Enhancement Network with a Public Test Dataset
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Over-sampling De-occlusion Attention Network for Prohibited Items Detection in Noisy X-ray Images
- Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation
- Multiple Object Tracking with Correlation Learning
- End-to-end Active Object Tracking and Its Real-world Deployment via Reinforcement Learning
- From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models
- Substitute Teacher Networks: Learning with Almost No Supervision
- Bag of Freebies for Training Object Detection Neural Networks
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- SurveilEdge: Real-time Video Query based on Collaborative Cloud-Edge Deep Learning
- Dynamic Region Division for Adaptive Learning Pedestrian Counting
- VarifocalNet: An IoU-aware Dense Object Detector
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Detection as Regression: Certified Object Detection by Median Smoothing
- EfficientDet: Scalable and Efficient Object Detection
- A Survey of Fish Tracking Techniques Based on Computer Vision
- Dynamic Adversarial Patch for Evading Object Detection Models
- Learning a Proposal Classifier for Multiple Object Tracking
- LLVIP: A Visible-infrared Paired Dataset for Low-light Vision
- Toward Intelligent Sensing: Intermediate Deep Feature Compression
- Pareto-Optimal Bit Allocation for Collaborative Intelligence
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- Vote from the Center: 6 DoF Pose Estimation in RGB-D Images by Radial Keypoint Voting
- Decision-based Universal Adversarial Attack
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Real-Time target detection in maritime scenarios based on YOLOv3 model
- Towards More Flexible and Accurate Object Tracking with Natural Language: Algorithms and Benchmark
- Non-local RoI for Cross-Object Perception
- Occluded Prohibited Items Detection: an X-ray Security Inspection Benchmark and De-occlusion Attention Module
- A Real-time Global Inference Network for One-stage Referring Expression Comprehension
- Multimodal Generative Models for Compositional Representation Learning
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Distilling Image Classifiers in Object Detectors
- FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
- Resource-Constrained Simultaneous Detection and Labeling of Objects in High-Resolution Satellite Images
- Scalable Image Coding for Humans and Machines
- A Survey on Sensor Technologies for Unmanned Ground Vehicles
- EDSL: An Encoder-Decoder Architecture with Symbol-Level Features for Printed Mathematical Expression Recognition
- IoU-aware Single-stage Object Detector for Accurate Localization
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis Network
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Detection and Tracking Meet Drones Challenge
- How Well Do Sparse Imagenet Models Transfer?
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- FVNet: 3D Front-View Proposal Generation for Real-Time Object Detection from Point Clouds
- WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
- General Instance Distillation for Object Detection
- A Practical Framework for ROI Detection in Medical Images -- a case study for hip detection in anteroposterior pelvic radiographs
- ARC: A Vision-based Automatic Retail Checkout System
- AI Pose Analysis and Kinematic Profiling of Range-of-Motion Variations in Resistance Training
- Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts
- AGE Challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
- Confidence Calibration with Bounded Error Using Transformations
- Monitoring Horses in Stalls: From Object to Event Detection
- Deep Knowledge Tracing with Learning Curves
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- A Modular Object Detection System for Humanoid Robots Using YOLO
- NarrationBot and InfoBot: A Hybrid System for Automated Video Description
- Detect Anything via Next Point Prediction
- Collaborative Multi-Robot Systems for Search and Rescue: Coordination and Perception
- A Modular AIoT Framework for Low-Latency Real-Time Robotic Teleoperation in Smart Cities
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Automated Defect Detection for Mass-Produced Electronic Components Based on YOLO Object Detection Models
- A Cost Effective Solution for Road Crack Inspection using Cameras and Deep Neural Networks
- Detection of Dataset Shifts in Learning-Enabled Cyber-Physical Systems using Variational Autoencoder for Regression
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- ThunderNet: Towards Real-time Generic Object Detection
- QANet -- Quality Assurance Network for Image Segmentation
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- Whole Body Model Predictive Control for Spin-Aware Quadrupedal Table Tennis
- Towards Adversarially Robust Object Detection
- Meta Adversarial Training against Universal Patches
- Video Surveillance for Road Traffic Monitoring
- Robust 2D/3D Vehicle Parsing in CVIS
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- EOLO: Embedded Object Segmentation only Look Once
- MaizeStandCounting (MaSC): Automated and Accurate Maize Stand Counting from UAV Imagery Using Image Processing and Deep Learning
- Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
- GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Investigating mixed traffic dynamics of pedestrians and non-motorized vehicles at urban intersections: Observation experiments and modelling
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Skin Lesion Classification Based on ResNet-50 Enhanced With Adaptive Spatial Feature Fusion
- Oriented RepPoints for Aerial Object Detection
- Humanly Certifying Superhuman Classifiers
- Using Images from a Video Game to Improve the Detection of Truck Axles
- YOLO-Based Defect Detection for Metal Sheets
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- Towards Palmprint Verification On Smartphones
- EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments
- Voxel-FPN: multi-scale voxel feature aggregation in 3D object detection from point clouds
- Machine Learning Based Analysis of Finnish World War II Photographers
- DetectorGuard: Provably Securing Object Detectors against Localized Patch Hiding Attacks
- Visual Relationship Detection using Scene Graphs: A Survey
- Scope Head for Accurate Localization in Object Detection
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Deep-RBF Networks Revisited: Robust Classification with Rejection
- Resolving Class Imbalance in Object Detection with Weighted Cross Entropy Losses
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
- MM-FSOD: Meta and metric integrated few-shot object detection
- NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination
- (AF)2-S3Net: Attentive Feature Fusion with Adaptive Feature Selection for Sparse Semantic Segmentation Network
- Dynamic DNN Decomposition for Lossless Synergistic Inference
- Camera Exposure Control for Robust Robot Vision with Noise-Aware Image Quality Assessment
- Representation Sharing for Fast Object Detector Search and Beyond
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- SOLO: A Simple Framework for Instance Segmentation
- A Survey of Autonomous Driving: Common Practices and Emerging Technologies
- IMMVP: An Efficient Daytime and Nighttime On-Road Object Detector
- Movement Tracks for the Automatic Detection of Fish Behavior in Videos
- An improved helmet detection method for YOLOv3 on an unbalanced dataset
- Learning Universal Shape Dictionary for Realtime Instance Segmentation
- Video Anomaly Detection by Estimating Likelihood of Representations
- A Comprehensive Review on Recent Methods and Challenges of Video Description
- Skeleon-Based Typing Style Learning For Person Identification
- Expedited Multi-Target Search with Guaranteed Performance via Multi-fidelity Gaussian Processes
- Assessing the Alignment of Popular CNNs to the Brain for Valence Appraisal
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- Automated Detecting and Placing Road Objects from Street-level Images
- Trustworthy Convolutional Neural Networks: A Gradient Penalized-based Approach
- Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high resolution RGB labelled images to develop and benchmark wheat head detection methods
- Enhanced Detection of Tiny Objects in Aerial Images
- Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- A MAC-less Neural Inference Processor Supporting Compressed, Variable Precision Weights
- Towards an efficient framework for Data Extraction from Chart Images
- Modification method for single-stage object detectors that allows to\n exploit the temporal behaviour of a scene to improve detection accuracy
- Video based real-time positional tracker
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Performance Optimization of YOLO-FEDER FusionNet for Robust Drone Detection in Visually Complex Environments
- A Survey of Deep Learning for Scientific Discovery
- A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
- DeepSperm: A robust and real-time bull sperm-cell detection in densely populated semen videos
- An Orientation Factor for Object-Oriented SLAM
- Trained Model Fusion for Object Detection using Gating Network
- Unifying Variational Inference and PAC-Bayes for Supervised Learning that Scales
- Automatically detecting pig position and posture by 2D camera imaging and deep learning
- CSIYOLO: An Intelligent CSI-based Scatter Sensing Framework for Integrated Sensing and Communication Systems
- Synthetic Occlusion Augmentation with Volumetric Heatmaps for the 2018 ECCV PoseTrack Challenge on 3D Human Pose Estimation
- Review: deep learning on 3D point clouds
- Rethinking Channel Dimensions for Efficient Model Design
- WebSight: A Vision-First Architecture for Robust Web Agents
- Locate then Segment: A Strong Pipeline for Referring Image Segmentation
- TKD: Temporal Knowledge Distillation for Active Perception
- Research on Fast Text Recognition Method for Financial Ticket Image
- Deep Learning Based FDD Non-Stationary Massive MIMO Downlink Channel Reconstruction
- VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Two-Stage Swarm Intelligence Ensemble Deep Transfer Learning (SI-EDTL) for Vehicle Detection Using Unmanned Aerial Vehicles
- DEPFusion: Dual-Domain Enhancement and Priority-Guided Mamba Fusion for UAV Multispectral Object Detection
- Parse Graph-Based Visual-Language Interaction for Human Pose Estimation
- FCOS: Fully Convolutional One-Stage Object Detection
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Deep learning visual analysis in laparoscopic surgery: a systematic review and diagnostic test accuracy meta-analysis
- Semantic Segmentation for Compound figures
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Detection of E-scooter Riders in Naturalistic Scenes
- FCOS: A simple and strong anchor-free object detector
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Plant diseases and pests detection based on deep learning: a review
- Video Analytics with Zero-streaming Cameras
- Assembly of randomly placed parts realized by using only one robot arm with a general parallel-jaw gripper
- Object Detection in Specific Traffic Scenes using YOLOv2
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
- Efficient Pipelines for Vision-Based Context Sensing
- Global Context Aware RCNN for Object Detection
- Non-local RoIs for Instance Segmentation
- An End-to-End Framework for Video Multi-Person Pose Estimation
- Empowering cyberphysical systems of systems with intelligence
- COVID-19 personal protective equipment detection using real-time deep learning methods
- Probabilistic two-stage detection
- QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in\n Object-oriented SLAM
- YOLOX: Exceeding YOLO Series in 2021
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Objects in Semantic Topology
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Referring Expression Comprehension: A Survey of Methods and Datasets
- Streamlining the Development of Active Learning Methods in Real-World Object Detection
- Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Weed Detection in Challenging Field Conditions: A Semi-Supervised Framework for Overcoming Shadow Bias and Data Scarcity
- Learning and Reasoning with the Graph Structure Representation in Robotic Surgery
- Exploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
- Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Deep Active Learning for Remote Sensing Object Detection
- Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
- Revisiting Knowledge Distillation for Object Detection
- CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
- FS-Net: Fast Shape-based Network for Category-Level 6D Object Pose Estimation with Decoupled Rotation Mechanism
- INSTA-YOLO: Real-Time Instance Segmentation
- G2L-Net: Global to Local Network for Real-time 6D Pose Estimation with Embedding Vector Features
- A Cost-Effective Framework for Predicting Parking Availability Using Geospatial Data and Machine Learning
- A system of vision sensor based deep neural networks for complex driving scene analysis in support of crash risk assessment and prevention
- Automatic lesion segmentation and Pathological Myopia classification in fundus images
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- Technical Report: Reactive Semantic Planning in Unexplored Semantic Environments Using Deep Perceptual Feedback
- A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network
- Semi-supervised Image Dehazing via Expectation-Maximization and Bidirectional Brownian Bridge Diffusion Models
- Colon Polyps Detection from Colonoscopy Images Using Deep Learning
- Lameness detection in dairy cows using pose estimation and bidirectional LSTMs
- Towards Powerful and Practical Patch Attacks for 2D Object Detection in Autonomous Driving
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Background Segmentation for Vehicle Re-Identification
- Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- v2e: From Video Frames to Realistic DVS Events
- CPGAN: Full-Spectrum Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Real-time landmark detection for precise endoscopic submucosal\n dissection via shape-aware relation network
- Non-anchor-based vehicle detection for traffic surveillance using bounding ellipses
- One-Shot Instance Segmentation
- Boosting Image Outpainting with Semantic Layout Prediction
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Strawberry Detection Using a Heterogeneous Multi-Processor Platform
- A Fast and Accurate One-Stage Approach to Visual Grounding
- Comparison-Based Convolutional Neural Networks for Cervical Cell/Clumps Detection in the Limited Data Scenario
- LF-YOLO: A Lighter and Faster YOLO for Weld Defect Detection of X-ray Image
- Domain Adaptive YOLO for One-Stage Cross-Domain Detection
- Real-Time Object Detection and Localization in Compressive Sensed Video on Embedded Hardware
- WQT and DG-YOLO: towards domain generalization in underwater object detection
- Decentralized Structural-RNN for Robot Crowd Navigation with Deep Reinforcement Learning
- WDR FACE: The First Database for Studying Face Detection in Wide Dynamic Range
- Fully Automatic Wound Segmentation with Deep Convolutional Neural Networks
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- LAMP: Large-Scale Autonomous Mapping and Positioning for Exploration of Perceptually-Degraded Subterranean Environments
- ULU: A Unified Activation Function
- YOLOv8-Based Deep Learning Model for Automated Poultry Disease Detection and Health Monitoring paper
- Benchmarking pig detection and tracking under diverse and challenging conditions
- Robust Processing-In-Memory Neural Networks via Noise-Aware Normalization
- CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps
- A Smartphone-based System for Real-time Early Childhood Caries Diagnosis
- Deep learning framework for crater detection and identification on the Moon and Mars
- Talking Detection In Collaborative Learning Environments
- Adversarial Attention Perturbations for Large Object Detection Transformers
- The Indirect Convolution Algorithm
- CoFF: Cooperative Spatial Feature Fusion for 3D Object Detection on Autonomous Vehicles
- Few-shot Object Detection with Self-adaptive Attention Network for Remote Sensing Images
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Privacy-Preserving Driver Drowsiness Detection with Spatial Self-Attention and Federated Learning
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- "Who is Driving around Me?" Unique Vehicle Instance Classification using Deep Neural Features
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image
- MultiResolution Attention Extractor for Small Object Detection
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- ContourRender: Detecting Arbitrary Contour Shape For Instance Segmentation In One Pass
- Object Recognition Datasets and Challenges: A Review
- Detection Method Based on Automatic Visual Shape Clustering for Pin-Missing Defect in Transmission Lines
- A Steel Surface Defect Detection Method Based on Lightweight Convolution Optimization
- Instance Segmentation with Point Supervision
- Gradient Harmonized Single-stage Detector
- Putting visual object recognition in context
- A2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Meaning guided video captioning
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- ABCD: Automatic Blood Cell Detection via Attention-Guided Improved YOLOX
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- Kidney Recognition in CT Using YOLOv3
- Egoshots, an ego-vision life-logging dataset and semantic fidelity metric to evaluate diversity in image captioning models
- Multi-adversarial Faster-RCNN for Unrestricted Object Detection
- DeepCore: Convolutional Neural Network for high pT jet tracking
- Optimizing Prediction Serving on Low-Latency Serverless Dataflow
- Towards Large Scale Geostatistical Methane Monitoring with Part-based Object Detection
- VASTA: A Vision and Language-assisted Smartphone Task Automation System
- Locating Cephalometric X-Ray Landmarks with Foveated Pyramid Attention
- Scalable Perception-Action-Communication Loops with Convolutional and Graph Neural Networks
- Focal and Global Knowledge Distillation for Detectors
- Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- Pathology-Aware Generative Adversarial Networks for Medical Image Augmentation
- Low-light Image Enhancement Algorithm Based on Retinex and Generative Adversarial Network
- Abnormal Behavior Detection Based on Target Analysis
- Semantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
- Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
- When We First Met: Visual-Inertial Person Localization for Co-Robot Rendezvous
- Intermediate Deep Feature Compression: the Next Battlefield of Intelligent Sensing
- SafeAccess+: An Intelligent System to make Smart Home Safer and Americans with Disability Act Compliant
- Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up Networks
- Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
- Flexible Vector Integration in Embedded RISC-V SoCs for End to End CNN Inference Acceleration
- N 2 C : Neural Network Controller Design Using Behavioral Cloning
- Design and Implementation of an Annotation-Driven Drone Autonomy Tool Using YOLOv8–V11 Architectures for Real-Time Object Detection and Distance Estimation
- Face Detection with Feature Pyramids and Landmarks
- Generating Multiple Objects at Spatially Distinct Locations
- Towards Universal Object Detection by Domain Attention
- Active Terahertz Imaging Dataset for Concealed Object Detection
- Synthetic-to-Real Domain Adaptation for Lane Detection
- OpenEI: An Open Framework for Edge Intelligence
- PP-YOLOv2: A Practical Object Detector
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection
- Dense Relation Distillation with Context-aware Aggregation for Few-Shot Object Detection
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Instance Scale Normalization for image understanding
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- An Inter-Layer Weight Prediction and Quantization for Deep Neural Networks based on a Smoothly Varying Weight Hypothesis
- Rethinking Natural Adversarial Examples for Classification Models
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- Phase Space Reconstruction Network for Lane Intrusion Action Recognition
- Analysis of Plant Nutrient Deficiencies Using Multi-Spectral Imaging and Optimized Segmentation Model
- Few-Shot Learning in Video and 3D Object Detection: A Survey
- Can 3D Adversarial Logos Cloak Humans?
- ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
- ShrinkBox: Backdoor Attack on Object Detection to Disrupt Collision Avoidance in Machine Learning-based Advanced Driver Assistance Systems
- Maximum Lifetime Analytics in IoT Networks
- Light-Weight RetinaNet for Object Detection
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
- SOD-YOLO: Enhancing YOLO-Based Detection of Small Objects in UAV Imagery
- Few-shot Object Detection with Feature Attention Highlight Module in Remote Sensing Images
- Universal Physical Camouflage Attacks on Object Detectors
- Selective Quantization Tuning for ONNX Models
- Report on UG2+ Challenge Track 1: Assessing Algorithms to Improve Video Object Detection and Classification from Unconstrained Mobility Platforms
- Semantic SLAM with Autonomous Object-Level Data Association
- Pixel Invisibility: Detecting Objects Invisible in Color Images
- MFPP: Morphological Fragmental Perturbation Pyramid for Black-Box Model Explanations
- High Speed and High Dynamic Range Video with an Event Camera
- Towards Autonomous Riding: A Review of Perception, Planning, and Control in Intelligent Two-Wheelers
- Rethinking of Radar's Role: A Camera-Radar Dataset and Systematic Annotator via Coordinate Alignment
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- Adversarial Robustness of Deep Convolutional Candlestick Learner
- Accurate and Robust Object-oriented SLAM with 3D Quadric Landmark Construction in Outdoor Environment
- TOG: Targeted Adversarial Objectness Gradient Attacks on Real-time Object Detection Systems
- Implementing AI-powered semantic character recognition in motor racing sports
- LiDAR Cluster First and Camera Inference Later: A New Perspective Towards Autonomous Driving
- DR-SPAAM: A Spatial-Attention and Auto-regressive Model for Person Detection in 2D Range Data
- Learning Spatial Fusion for Single-Shot Object Detection
- ACR-Pose: Adversarial Canonical Representation Reconstruction Network for Category Level 6D Object Pose Estimation
- Salience Biased Loss for Object Detection in Aerial Images
- Few-shot Object Detection via Feature Reweighting
- OSKDet: Towards Orientation-sensitive Keypoint Localization for Rotated Object Detection
- Visibility Guided NMS: Efficient Boosting of Amodal Object Detection in Crowded Traffic Scenes
- Robotic Grasp Manipulation Using Evolutionary Computing and Deep Reinforcement Learning
- One Patch to Rule Them All: Transforming Static Patches into Dynamic Attacks in the Physical World
- Universal Adversarial Perturbations: A Survey
- Residual Squeeze-and-Excitation Network for Fast Image Deraining
- UWGAN: Underwater GAN for Real-world Underwater Color Restoration and Dehazing
- Boosting ship detection in SAR images with complementary pretraining techniques
- Does Thermal data make the detection systems more reliable?
- Object-aware Feature Aggregation for Video Object Detection
- BAF-Detector: An Efficient CNN-Based Detector for Photovoltaic Cell Defect Detection
- Towards navigation without precise localization: Weakly supervised learning of goal-directed navigation cost map
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Multi-regime analysis for computer vision-based traffic surveillance using a change-point detection algorithm
- PPGN: Phrase-Guided Proposal Generation Network For Referring Expression Comprehension
- Dense Multiscale Feature Fusion Pyramid Networks for Object Detection in UAV-Captured Images
- RGB Stream Is Enough for Temporal Action Detection
- A Lightweight and Robust Framework for Real-Time Colorectal Polyp Detection Using LOF-Based Preprocessing and YOLO-v11n
- Real-Time Text Detection and Recognition
- Compact 3D Map-Based Monocular Localization Using Semantic Edge Alignment
- MultiScope: Efficient Video Pre-processing for Exploratory Video Analytics
- 3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
- EyeSeg: An Uncertainty-Aware Eye Segmentation Framework for AR/VR
- When Schrödinger Bridge Meets Real-World Image Dehazing with Unpaired Training
- Channel-wise Alignment for Adaptive Object Detection
- A Deep Learning-Based Autonomous RobotManipulator for Sorting Application
- Stereo-based 3D Anomaly Object Detection for Autonomous Driving: A New Dataset and Baseline
- DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
- From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
- Cross-Resolution SAR Target Detection Using Structural Hierarchy Adaptation and Reliable Adjacency Alignment
- Image-based phenotyping of diverse Rice (Oryza Sativa L.) Genotypes
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- YOLO-LITE: A Real-Time Object Detection Algorithm Optimized for Non-GPU Computers
- DeepApple: Deep Learning-based Apple Detection using a Suppression Mask R-CNN
- A Graph Attention Spatio-temporal Convolutional Network for 3D Human Pose Estimation in Video
- The Amazing Race TM: Robot Edition
- YOLO-APD: Enhancing YOLOv8 for Robust Pedestrian Detection on Complex Road Geometries
- Domain Adaptive Object Detection via Feature Separation and Alignment
- Visual Identification of Individual Holstein-Friesian Cattle via Deep Metric Learning
- From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
- Deep learning-based hierarchical cattle behavior recognition with spatio-temporal information
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- Know Your Surroundings: Panoramic Multi-Object Tracking by Multimodality Collaboration
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- 2.5D Object Detection for Intelligent Roadside Infrastructure
- Computer Vision-based Social Distancing Surveillance Solution with Optional Automated Camera Calibration for Large Scale Deployment
- R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous Driving
- An Efficient UAV-based Artificial Intelligence Framework for Real-Time Visual Tasks
- Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
- Red grape detection with accelerated artificial neural networks in the FPGA's programmable logic
- A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
- The Problem of Fragmented Occlusion in Object Detection
- Robust Real-Time Pedestrian Detection on Embedded Devices
- Understanding Trade offs When Conditioning Synthetic Data
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- Dynamic Refinement Network for Oriented and Densely Packed Object Detection
- All-In-One Underwater Image Enhancement using Domain-Adversarial Learning
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
- Adsorbate-induced formation of a surface-polarity-driven nonperiodic superstructure
- UMLE: Unsupervised Multi-discriminator Network for Low Light Enhancement
- Underwater Fish Detection using Deep Learning for Water Power Applications
- A deep learning-based framework for an automated defect detection system for sewer pipes
- Automatic detection of sewer defects based on improved you only look once algorithm
- X-LineNet: Detecting Aircraft in Remote Sensing Images by a pair of Intersecting Line Segments
- PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
- Semi-Supervised Surface Anomaly Detection of Composite Wind Turbine Blades From Drone Imagery
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- Congestion Analysis of Convolutional Neural Network-Based Pedestrian Counting Methods on Helicopter Footage
- Can WiFi Estimate Person Pose?
- Variational Pedestrian Detection
- Ink Marker Segmentation in Histopathology Images Using Deep Learning
- Improving Token-based Object Detection with Video
- Testbed for Connected Artificial Intelligence using Unmanned Aerial Vehicles and Convolutional Pose Machines
- Understanding Uncertainty of Edge Computing: New Principle and Design Approach
- Energy Drain of the Object Detection Processing Pipeline for Mobile Devices: Analysis and Implications
- Deep RNN Framework for Visual Sequential Applications
- YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
- Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
- Toward Autonomous Rotation-Aware Unmanned Aerial Grasping
- Learning to adapt class-specific features across domains for semantic segmentation
- LogoDet-3K: A Large-Scale Image Dataset for Logo Detection
- Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- CBAM-STN-TPS-YOLO: Enhancing Agricultural Object Detection through Spatially Adaptive Attention Mechanisms
- STA: Adversarial Attacks on Siamese Trackers
- Multi-Modal Zero-Shot Sign Language Recognition
- Exploring 2D Data Augmentation for 3D Monocular Object Detection
- Pattern-Based Phase-Separation of Tracer and Dispersed Phase Particles in Two-Phase Defocusing Particle Tracking Velocimetry
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- PatternMonitor: a whole pipeline with a much higher level of automation for guessing Android lock pattern based on videos
- LunarLoc: Segment-Based Global Localization on the Moon
- A keypoint-based method for detecting weed growth points in corn field environments
- Leveraging CNN and IoT for Effective E-Waste Management
- Detecting and Matching Related Objects with One Proposal Multiple Predictions
- You Cannot Easily Catch Me: A Low-Detectable Adversarial Patch for Object Detectors
- Learning to Generate Diverse Dance Motions with Transformer
- Multi-frame Collaboration for Effective Endoscopic Video Polyp Detection via Spatial-Temporal Feature Transformation
- Interpretable and Granular Video-Based Quantification of Motor Characteristics from the Finger Tapping Test in Parkinson Disease
- Underwater object detection using Invert Multi-Class Adaboost with deep learning
- PointINS: Point-based Instance Segmentation
- The Detection of Thoracic Abnormalities ChestX-Det10 Challenge Results
- YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework
- Hallucination Improves Few-Shot Object Detection
- Robustness Enhancement of Object Detection in Advanced Driver Assistance Systems (ADAS)
- Self-supervised Low Light Image Enhancement and Denoising
- Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- ADAM-Dehaze: Adaptive Density-Aware Multi-Stage Dehazing for Improved Object Detection in Foggy Conditions
- CSI2Image: Image Reconstruction from Channel State Information Using Generative Adversarial Networks
- Bring Your Own Codegen to Deep Learning Compiler
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- Object Detection-Based Variable Quantization Processing
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Improving One-stage Visual Grounding by Recursive Sub-query Construction
- Good Practices and A Strong Baseline for Traffic Anomaly Detection
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- Simultaneous Learning from Human Pose and Object Cues for Real-Time Activity Recognition
- Multi-Scale Boosted Dehazing Network with Dense Feature Fusion
- Sequence Level Semantics Aggregation for Video Object Detection
- On the Natural Robustness of Vision-Language Models Against Visual Perception Attacks in Autonomous Driving
- YOLO5Face: Why Reinventing a Face Detector
- UG2+ Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments
- DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
- Efficiency Robustness of Dynamic Deep Learning Systems
- On Designing Computing Systems for Autonomous Vehicles: a PerceptIn Case Study
- Graph Neural Networks for Natural Language Processing: A Survey
- Spatial Process Mining
- M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network
- On Applying Machine Learning/Object Detection Models for Analysing Digitally Captured Physical Prototypes from Engineering Design Projects
- Few-Shot Learning for Road Object Detection
- Multiple Object Tracking with Motion and Appearance Cues
- Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
- Image Conditioned Keyframe-Based Video Summarization Using Object Detection
- Improvement of Classification in One-Stage Detector
- Scalable Aerial GNSS Localization for Marine Robots
- FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review
- 360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
- Joint-SRVDNet: Joint Super Resolution and Vehicle Detection Network
- RAPiD: Rotation-Aware People Detection in Overhead Fisheye Images
- NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
- Scale-Aware Trident Networks for Object Detection
- Adversarial Examples in Modern Machine Learning: A Review
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- TagMe: GPS-Assisted Automatic Object Annotation in Videos
- Towards Better Object Detection in Scale Variation with Adaptive Feature Selection
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- Estimation of Closest In-Path Vehicle (CIPV) by Low-Channel LiDAR and Camera Sensor Fusion for Autonomous Vehicle
- MambaNeXt-YOLO: A Hybrid State Space Model for Real-time Object Detection
- DiagNet: Detecting Objects using Diagonal Constraints on Adjacency Matrix of Graph Neural Network
- Adaptive Object Detection with Dual Multi-Label Prediction
- Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector
- Deeply Activated Salient Region for Instance Search
- Ice Hockey Puck Localization Using Contextual Cues
- A New Action Recognition Framework for Video Highlights Summarization in Sporting Events
- Drowsiness Detection Based On Driver Temporal Behavior Using a New Developed Dataset
- Fooling Detection Alone is Not Enough: First Adversarial Attack against Multiple Object Tracking
- A Dynamic Transformer Network for Vehicle Detection
- Auto-calibration Method Using Stop Signs for Urban Autonomous Driving Applications
- Real-time Human-Robot Collaborative Manipulations of Cylindrical and Cubic Objects via Geometric Primitives and Depth Information
- Multi-person 3D Pose Estimation in Crowded Scenes Based on Multi-View Geometry
- DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better
- VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment
- Object detection on aerial imagery using CenterNet
- ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos
- ChannelExplorer: Exploring Class Separability Through Activation Channel Visualization
- MIXER: A Principled Framework for Multimodal, Multiway Data Association
- Pseudo-Labeling Driven Refinement of Benchmark Object Detection Datasets via Analysis of Learning Patterns
- Probabilistic Oriented Object Detection in Automotive Radar
- Metamorphic Testing for Object Detection Systems
- FCOSR: A Simple Anchor-free Rotated Detector for Aerial Object Detection
- Fooling thermal infrared pedestrian detectors in real world using small bulbs
- AWML: An Open-Source ML-based Robotics Perception Framework to Deploy for ROS-based Autonomous Driving Software
- Test-time Vocabulary Adaptation for Language-driven Object Detection
- Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural Accelerators
- 2nd Place Solution for Waymo Open Dataset Challenge -- Real-time 2D Object Detection
- Multi-Task Incremental Learning for Object Detection
- DiG-Net: Enhancing Quality of Life through Hyper-Range Dynamic Gesture Recognition in Assistive Robotics
- Discriminative Semantic Feature Pyramid Network with Guided Anchoring for Logo Detection
- One-Shot Imitation Filming of Human Motion Videos
- Spectro-Temporal RF Identification using Deep Learning
- Efficient Pig Counting in Crowds with Keypoints Tracking and Spatial-aware Temporal Response Filtering
- Modulating Localization and Classification for Harmonized Object Detection
- A Robot-Assisted Approach to Small Talk Training for Adults with ASD
- TinaFace: Strong but Simple Baseline for Face Detection
- MuSe 2020 -- The First International Multimodal Sentiment Analysis in Real-life Media Challenge and Workshop
- A Comparative Study on Effects of Original and Pseudo Labels for Weakly Supervised Learning for Car Localization Problem
- Domain-Specific Suppression for Adaptive Object Detection
- MSFNet-CPD: Multi-Scale Cross-Modal Fusion Network for Crop Pest Detection
- Semantic Feature Matching for Robust Mapping in Agriculture
- OrcVIO: Object residual constrained Visual-Inertial Odometry
- Composition and Configuration Patterns in Multiple-View Visualizations
- RTFN: A Robust Temporal Feature Network for Time Series Classification
- AmphibianDetector: adaptive computation for moving objects detection
- YOLO-SPCI: Enhancing Remote Sensing Object Detection via Selective-Perspective-Class Integration
- VisAlgae 2023: A Dataset and Challenge for Algae Detection in Microscopy Images
- Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
- Classification of oil palm tree conditions from UAV imagery using the YOLO object detector
- Defense-guided Transferable Adversarial Attacks
- Pedestrian Detection in 3D Point Clouds using Deep Neural Networks
- Single-Stage 6D Object Pose Estimation
- Learning a Domain Classifier Bank for Unsupervised Adaptive Object Detection
- CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection
- MeshAdv: Adversarial Meshes for Visual Recognition
- Deep Submodular Networks for Extractive Data Summarization
- Object grasping planning for the situation when soft and rigid objects are mixed together
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
- Autocamera Calibration for traffic surveillance cameras with wide angle lenses
- Leveraging Bottom-Up and Top-Down Attention for Few-Shot Object Detection
- Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
- Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
- Can We Automate Diagrammatic Reasoning?
- CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
- InfoDet: A Dataset for Infographic Element Detection
- A Software Architecture for Autonomous Vehicles: Team LRM-B Entry in the First CARLA Autonomous Driving Challenge
- MSDU-net: A Multi-Scale Dilated U-net for Blur Detection
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- Extending Dataset Pruning to Object Detection: A Variance-based Approach
- Enhancing Cross-task Black-Box Transferability of Adversarial Examples with Dispersion Reduction
- Detecting 11K Classes: Large Scale Object Detection without Fine-Grained Bounding Boxes
- AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
- AdvReal: Physical Adversarial Patch Generation Framework for Security Evaluation of Object Detection Systems
- Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining
- Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization
- Embodied Visual Recognition
- BadSR: Stealthy Label Backdoor Attacks on Image Super-Resolution
- SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
- Rate-Accuracy Bounds in Visual Coding for Machines
- TE-YOLOF: Tiny and efficient YOLOF for blood cell detection
- Depth Based Semantic Scene Completion with Position Importance Aware Loss
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- Demiguise Attack: Crafting Invisible Semantic Adversarial Perturbations with Perceptual Similarity
- RepPoints: Point Set Representation for Object Detection
- Learning More with Less: GAN-based Medical Image Augmentation
- FPAENet: Pneumonia Detection Network Based on Feature Pyramid Attention Enhancement
- Segmentation-driven 6D Object Pose Estimation
- Enlisting 3D Crop Models and GANs for More Data Efficient and Generalizable Fruit Detection
- Rethinking Features-Fused-Pyramid-Neck for Object Detection
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Additive Noise Annealing and Approximation Properties of Quantized Neural Networks
- An Overview of Arithmetic Adaptations for Inference of Convolutional Neural Networks on Re-configurable Hardware
- AGI-Elo: How Far Are We From Mastering A Task?
- DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object Detection
- Localize to Classify and Classify to Localize: Mutual Guidance in Object Detection
- Residual Error: a New Performance Measure for Adversarial Robustness
- Development of a conversing and body temperature scanning autonomously navigating robot to help screen for COVID-19
- Towards in-store multi-person tracking using head detection and track heatmaps
- Adversarial Semantic Contour for Object Detection
- Depth-conditioned Dynamic Message Propagation for Monocular 3D Object Detection
- A Serverless Cloud-Fog Platform for DNN-Based Video Analytics with Incremental Learning
- DAF-NET: a saliency based weakly supervised method of dual attention fusion for fine-grained image classification
- Stochastic-YOLO: Efficient Probabilistic Object Detection under Dataset Shifts
- Mask R-CNN Based Object Detection for Intelligent Wireless Power Transfer
- SIMPL: Generating Synthetic Overhead Imagery to Address Zero-shot and Few-Shot Detection Problems
- Self-supervised Robust Object Detectors from Partially Labelled Datasets
- Driving Datasets Literature Review
- Multi-directional Bicycle Robot for Steel Structure Inspection
- MLMA-Net: multi-level multi-attentional learning for multi-label object detection in textile defect images
- Thermal Detection of People with Mobility Restrictions for Barrier Reduction at Traffic Lights Controlled Intersections
- OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement
- CenterFace: Joint Face Detection and Alignment Using Face as Point
- Rearchitecting Classification Frameworks For Increased Robustness
- FSD: Feature Skyscraper Detector for Stem End and Blossom End of Navel Orange
- The distance between the weights of the neural network is meaningful
- Airplane Detection Based on Mask Region Convolution Neural Network
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- Shared Mobile-Cloud Inference for Collaborative Intelligence
- Custom Object Detection via Multi-Camera Self-Supervised Learning
- LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
- A Study on Trees's Knots Prediction from their Bark Outer-Shape
- Domain-specific AI segmentation of IMPDH2 rod/ring structures in mouse embryonic stem cells
- Underwater object detection in sonar imagery with detection transformer and Zero-shot neural architecture search
- The Application of Deep Learning for Lymph Node Segmentation: A Systematic Review
- Uncertainty-Aware Voxel based 3D Object Detection and Tracking with von-Mises Loss
- Progressive Localization Networks for Language-based Moment Localization
- Improving the Transferability of Adversarial Examples with New Iteration Framework and Input Dropout
- Anchors Based Method for Fingertips Position Estimation from a Monocular RGB Image using Deep Neural Network
- NAS-FCOS: Efficient Search for Object Detection Architectures
- White Light Specular Reflection Data Augmentation for Deep Learning Polyp Detection
- Road Damage Detection and Classification with Detectron2 and Faster R-CNN
- OODTE: A Differential Testing Engine for the ONNX Optimizer
- Adversarial Robustness of Deep Learning Models for Inland Water Body Segmentation from SAR Images
- Localization Distillation for Dense Object Detection
- Complex-Object Visual Inspection via Multiple Lighting Configurations
- Integrating NVIDIA Deep Learning Accelerator (NVDLA) with RISC-V SoC on FireSim
- Fast Object Detection with Latticed Multi-Scale Feature Fusion
- Faster object tracking pipeline for real time tracking
- Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
- Spatiotemporal Action Recognition in Restaurant Videos
- Depth-Aware Action Recognition: Pose-Motion Encoding through Temporal Heatmaps
- ST-DETR: Spatio-Temporal Object Traces Attention Detection Transformer
- Knowledge Distillation By Sparse Representation Matching
- AGSFCOS: Based on attention mechanism and Scale-Equalizing pyramid network of object detection
- Visual Trajectory Prediction of Vessels for Inland Navigation
- A Possible Reason for why Data-Driven Beats Theory-Driven Computer Vision
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- Teaching Perception
- RMOPP: Robust Multi-Objective Post-Processing for Effective Object Detection
- Live Reconstruction of Large-Scale Dynamic Outdoor Worlds
- Zero-Shot Scene Graph Relation Prediction through Commonsense Knowledge Integration
- Adaptive Protein Tokenization
- Dynamic and Static Object Detection Considering Fusion Regions and Point-wise Features
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Classic versus deep learning approaches to address computer vision challenges
- Real-World Image Datasets for Federated Learning
- Target Reaching Behaviour for Unfreezing the Robot in a Semi-Static and Crowded Environment
- More Reliable AI Solution: Breast Ultrasound Diagnosis Using Multi-AI Combination
- Automatic Mapping with Obstacle Identification for Indoor Human Mobility Assessment
- Large-scale Gastric Cancer Screening and Localization Using Multi-task Deep Neural Network
- Research on Optimization Method of Multi-scale Fish Target Fast Detection Network
- A Unified Optimization Approach for CNN Model Inference on Integrated GPUs
- EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- Category-Level and Open-Set Object Pose Estimation for Robotics
- ODExAI: A Comprehensive Object Detection Explainable AI Evaluation
- Transcending Dimensions using Generative AI: Real-Time 3D Model Generation in Augmented Reality
- Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- PerfCam: Digital Twinning for Production Lines Using 3D Gaussian Splatting and Vision Models
- Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link Model
- WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects
- Semantic Reality: Interactive Context-Aware Visualization of Inter-Object Relationships in Augmented Reality
- Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference
- Did you miss it? Automatic lung nodule detection combined with gaze information improves radiologists' screening performance
- Ground-aware Monocular 3D Object Detection for Autonomous Driving
- A Decade of You Only Look Once (YOLO) for Object Detection: A Review
- Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
- Wide-Depth-Range 6D Object Pose Estimation in Space
- A Bioinspired Retinal Neural Network for Accurately Extracting Small-Target Motion Information in Cluttered Backgrounds
- Cooperating RPN's Improve Few-Shot Object Detection
- ABCP: Automatic Block-wise and Channel-wise Network Pruning via Joint Search
- Compressed Object Detection
- Generative Fields: Uncovering Hierarchical Feature Control for StyleGAN via Inverted Receptive Fields
- YOLOv14:Unified Cross-Domain Real-Time Object Detectionwith Adaptive Multi-View Representation
- Stampnet: Unsupervised Multi-Class Object Discovery
- A Smart, Efficient, and Reliable Parking Surveillance System With Edge Artificial Intelligence on IoT Devices
- SMOT: Single-Shot Multi Object Tracking
- Enabling Incremental Knowledge Transfer for Object Detection at the Edge
- MMFNet: A Multi-modality MRI Fusion Network for Segmentation of Nasopharyngeal Carcinoma
- The Cube++ Illumination Estimation Dataset
- Intelligent Defect Detection Method for Additive Manufactured Lattice Structures Based on a Modified YOLOv3 Model
- A Generic Object Re-identification System for Short Videos
- Cell Detection from Imperfect Annotation by Pseudo Label Selection Using P-classification
- Grounding-Tracking-Integration
- A survey of deep learning techniques for autonomous driving
- A Unified Framework for Generic, Query-Focused, Privacy Preserving and Update Summarization using Submodular Information Measures
- A Computer Vision Based Beamforming Scheme for Millimeter Wave Communication in LOS Scenarios
- FRDet: Balanced and Lightweight Object Detector based on Fire-Residual Modules for Embedded Processor of Autonomous Driving
- SA-Net: Robust State-Action Recognition for Learning from Observations
- DuBox: No-Prior Box Objection Detection via Residual Dual Scale Detectors
- KORSAL: Key-point Detection based Online Real-Time Spatio-Temporal\n Action Localization
- DeepTEGINN: Deep Learning Based Tools to Extract Graphs from Images of Neural Networks
- Pyramid Vector Quantization and Bit Level Sparsity in Weights for Efficient Neural Networks Inference
- AFD-Net: Adaptive Fully-Dual Network for Few-Shot Object Detection
- VISTA-OCR: Towards generative and interactive end to end OCR models
- Automatic Generation of Machine Learning Synthetic Data Using ROS
- Reactive Human-to-Robot Handovers of Arbitrary Objects
- Enhance Then Search: An Augmentation-Search Strategy with Foundation Models for Cross-Domain Few-Shot Object Detection
- Affinity LCFCN: Learning to Segment Fish with Weak Supervision
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- R-AGNO-RPN: A LIDAR-Camera Region Deep Network for Resolution-Agnostic Detection
- A Deep Learning-Based FPGA Function Block Detection Method With Bitstream to Image Transformation
- D2-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios
- Object detection for graphical user interface: old fashioned or deep learning or a combination?
- Learning on Hardware: A Tutorial on Neural Network Accelerators and Co-Processors
- Segmentation-Based Bounding Box Generation for Omnidirectional Pedestrian Detection
- HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
- Autonomous and Cooperative Design of the Monitor Positions for a Team of UAVs to Maximize the Quantity and Quality of Detected Objects
- Deep Sensing of Urban Waterlogging
- Detecting The Objects on The Road Using Modular Lightweight Network
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- LIFT-SLAM: A deep-learning feature-based monocular visual SLAM method
- LittleYOLO-SPP: A delicate real-time vehicle detection algorithm
- Towards Automated Swimming Analytics Using Deep Neural Networks
- Multiple Myeloma Cancer Cell Instance Segmentation
- Vehicle detection and counting from VHR satellite images: efforts and open issues
- Six-channel Image Representation for Cross-domain Object Detection
- NENET: An Edge Learnable Network for Link Prediction in Scene Text
- Generative adversarial network with object detector discriminator for enhanced defect detection on ultrasonic B-scans
- Learn an Effective Lip Reading Model without Pains
- Behavioural Pattern Discovery from Collections of Egocentric Photo-Streams
- Semantic Relation Preserving Knowledge Distillation for Image-to-Image Translation
- Structure-aware completion of photogrammetric meshes in urban road environment
- Accurate Tracking of Arabidopsis Root Cortex Cell Nuclei in 3D Time-Lapse Microscopy Images Based on Genetic Algorithm
- A Review of YOLOv12: Attention-Based Enhancements vs. Previous Versions
- Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection
- MEC-Patch: Visible-Infrared Cross-Modal Adversarial Attack Driven by Intrinsic Material Emissivity Laws
- Pose-aware instance segmentation framework from cone beam CT images for tooth segmentation
- Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAVTarget Detection
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
- EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection
- CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared Detectors
- Balancing Two Classifiers via A Simplex ETF Structure for Model Calibration
- Small Object Detection with YOLO: A Performance Analysis Across Model Versions and Hardware
- Matching Representations of Explainable Artificial Intelligence and Eye Gaze for Human-Machine Interaction
- PapMOT: Exploring Adversarial Patch Attack against Multiple Object Tracking
- Light-YOLOv8-Flame: A Lightweight High-Performance Flame Detection Algorithm
- Adversarial Examples in Environment Perception for Automated Driving (Review)
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Empirical Upper Bound, Error Diagnosis and Invariance Analysis of Modern Object Detectors
- Strawberry Detection using Mixed Training on Simulated and Real Data
- iTV: Inferring Traffic Violation-Prone Locations With Vehicle Trajectories and Road Environment Data
- Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax
- You Only Look Once [wikipedia]
- Small object detection [wikipedia]
Discussions
- Not an essay but I think about this section from the YOLOv3 paper often https://arxiv.org/pdf/1804.02767 [bsky, 74 points, 3 comments]
- how did i miss this paper?? arxiv.org/abs/1804.02767 [bsky, 28 points, 5 comments]
- I’ve read perhaps dozens of CV research papers and found that the most important paragraph, by a wide margin, to be the conclusion of the yoloV3 tech report… arxiv.org/abs/1804.02767 [bsky, 11 points, 0 comments]
- My award for most honest paper goes to: arxiv.org/abs/1804.02767 An arxiv-only paper describing v3 of YOLO, a vision algo that already had two published papers. With huge adoption already, authors wer [bsky, 4 points, 0 comments]
- YOLOv3: An Incremental Improvement [hn, 1 points, 0 comments]
- Today a colleague forwarded this article to me, from Joseph Redmon and Ali Farhadi YOLOv3: An Incremental Improvement arxiv.org/abs/1804.02767 The relevant passage: [bsky, 1 points, 0 comments]
- Some academic writing style palate cleanser for frustrated paper writers: arxiv.org/pdf/1804.02767 [bsky, 1 points, 0 comments]
- Reminded today that this paper exists: arxiv.org/pdf/1804.02767 [bsky, 1 points, 0 comments]
- This guy arxiv.org/abs/1804.02767 [bsky, 1 points, 0 comments]
Related