Rich feature hierarchies for accurate object detection and semantic\n segmentation
2013/11/11 by Ross Girshick, Jeff Donahue, Girshick, Ross +5 · 242 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1311.2524
openalex publication_date 2013/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Object detection performance, as measured on the canonical PASCAL VOC\ndataset, has plateaued in the last few years. The best-performing methods are\ncomplex ensemble systems that typically combine multiple low-level image\nfeatures with high-level context. In this paper, we propose a simple and\nscalable detection algorithm that improves mean average precision (mAP) by more\nthan 30% relative to the previous best result on VOC 2012---achieving a mAP of\n53.3%. Our approach combines two key insights: (1) one can apply high-capacity\nconvolutional neural networks (CNNs) to bottom-up region proposals in order to\nlocalize and segment objects and (2) when labeled training data is scarce,\nsupervised pre-training for an auxiliary task, followed by domain-specific\nfine-tuning, yields a significant performance boost. Since we combine region\nproposals with CNNs, we call our method R-CNN: Regions with CNN features. We\nalso compare R-CNN to OverFeat, a recently proposed sliding-window detector\nbased on a similar CNN architecture. We find that R-CNN outperforms OverFeat by\na large margin on the 200-class ILSVRC2013 detection dataset. Source code for\nthe complete system is available at http://www.cs.berkeley.edu/~rbg/rcnn.\n
Citations
Cited by
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Towards High-Level Semantic Intelligence
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- FlowDet: Unifying Object Detection and Generative Transport Flows
- SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal
- Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
- Reducing Label Dependency in Human Activity Recognition with Wearables: From Supervised Learning to Novel Weakly Self-Supervised Approaches
- Reliable Detection of Minute Targets in High-Resolution Aerial Imagery across Temporal Shifts
- Optimal transport unlocks end-to-end learning for single-molecule localization
- YOLO and SGBM Integration for Autonomous Tree Branch Detection and Depth Estimation in Radiata Pine Pruning Applications
- BlinkBud: Detecting Hazards from Behind via Sampled Monocular 3D Detection on a Single Earbud
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- Intelligent Image Search Algorithms Fusing Visual Large Models
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- LMSeg: an end-to-end geometric message-passing network on barycentric dualgraphs for large-scale landscape mesh segmentation
- Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- A lightweight detector for real-time detection of remote sensing images
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection
- MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
- Fast Data Attribution for Text-to-Image Models
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- Deep Roto-Translation Scattering for Object Classification
- Neural Network Interoperability Across Platforms
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- A Hybrid YOLOv5-SSD IoT-Based Animal Detection System for Durian Plantation Protection
- Sewer pipeline condition assessment and defect detection using computer vision
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- An Efficient and Generalizable Transfer Learning Method for Weather Condition Detection on Ground Terminals
- SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- Doubly Smoothed Density Estimation with Application on Miners' Unsafe Act Detection
- WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- Learning Human-Object Interaction as Groups
- Through the Lens of Doubt: Robust and Efficient Uncertainty Estimation for Visual Place Recognition
- Document Intelligence in the Era of Large Language Models: A Survey
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Explaining raw data complexity to improve satellite onboard processing
- Referring Expression Comprehension for Small Objects
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Motion-Aware Transformer for Multi-Object Tracking
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- DENet: Dual-Path Edge Network with Global-Local Attention for Infrared Small Target Detection
- Learning Detection with Diverse Proposals
- Fine-grained Categorization and Dataset Bootstrapping using Deep Metric Learning with Humans in the Loop
- Learning Convolutional Networks for Content-weighted Image Compression
- Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs\n and Bilateral Neural Networks
- Parsing Occluded People by Flexible Compositions
- Learning Multi-Domain Convolutional Neural Networks for Visual Tracking
- Convolutional Channel Features
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to\n Mask Weights
- Face Detection through Scale-Friendly Deep Convolutional Networks
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Learning to Segment Every Thing
- HitoMi-Cam: A Shape-Agnostic Person Detection Method Using the Spectral Characteristics of Clothing
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models
- Incorporating the Refractory Period into Spiking Neural Networks through Spike-Triggered Threshold Dynamics
- Weakly Supervised Object Localization Using Things and Stuff Transfer
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Enhanced Detection of Tiny Objects in Aerial Images
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- Saccadic Vision for Fine-Grained Visual Classification
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Region-Aware Deformable Convolutions
- A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
- Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals
- Attend in groups: a weakly-supervised deep learning framework for learning from web data
- Deep Multi-camera People Detection
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- DeepMVS: Learning Multi-view Stereopsis
- Two-Stage Swarm Intelligence Ensemble Deep Transfer Learning (SI-EDTL) for Vehicle Detection Using Unmanned Aerial Vehicles
- Fast-SCNN: Fast Semantic Segmentation Network
- Deep Image Retrieval: Learning global representations for image search
- P3-SAM: Native 3D Part Segmentation
- Analysing domain shift factors between videos and images for object\n detection
- Pothole Detection and Recognition based on Transfer Learning
- A New Hybrid Model of Generative Adversarial Network and You Only Look Once Algorithm for Automatic License-Plate Recognition
- Towards Open World Detection: A Survey
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- Efficient Odd-One-Out Anomaly Detection
- A Convolutional Hierarchical Deep-learning Neural Network (C-HiDeNN) Framework for Non-linear Finite Element Analysis
- A flexible FPGA accelerator for convolutional neural networks
- The Role of Context Selection in Object Detection
- Non-local RoIs for Instance Segmentation
- Perceptual Generative Adversarial Networks for Small Object Detection
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and\n Recurrent Neural Networks
- Data Distillation: Towards Omni-Supervised Learning
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep\n Architecture
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Amodal Completion and Size Constancy in Natural Scenes
- Spot the Difference by Object Detection
- Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
- Dice Loss for Data-imbalanced NLP Tasks
- Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection
- A Pursuit of Temporal Accuracy in General Activity Detection
- D2-Mamba: Dual-Scale Fusion and Dual-Path Scanning with SSMs for Shadow Removal
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- Data Shift of Object Detection in Autonomous Driving
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- Recurrent Residual Module for Fast Inference in Videos
- Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
- Generic Tubelet Proposals for Action Localization
- Joint Object and Part Segmentation using Deep Learned Potentials
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- Context Encoding for Semantic Segmentation
- Monocular Object Instance Segmentation and Depth Ordering with CNNs
- Deep learning for class-generic object detection
- Deep Learning for Single-View Instance Recognition
- Benchmarking pig detection and tracking under diverse and challenging conditions
- What Holds Back Open-Vocabulary Segmentation?
- CityPersons: A Diverse Dataset for Pedestrian Detection
- Unsupervised Learning of Visual Representations using Videos
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Contextual Action Recognition with R*CNN
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- IAUNet: Instance-Aware U-Net
- RMT-PPAD: Real-time Multi-task Learning for Panoptic Perception in Autonomous Driving
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Backdoor Attacks on Deep Learning Face Detection
- FusionNet: 3D Object Classification Using Multiple Data Representations
- Improved Person Detection on Omnidirectional Images with Non-maxima Suppression
- Tricks and Plug-ins for Gradient Boosting in Image Classification
- Multi-scale Location-aware Kernel Representation for Object Detection
- Object Recognition Datasets and Challenges: A Review
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- SARD: A YOLOv8-Based System for Solar Active Region Detection with SDO/HMI Magnetograms
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Class Subset Selection for Transfer Learning using Submodularity
- MITOS-RCNN: A Novel Approach to Mitotic Figure Detection in Breast Cancer Histopathology Images using Region Based Convolutional Neural Networks
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- JDATT: A Joint Distillation Framework for Atmospheric Turbulence Mitigation and Target Detection
- LCNN: Lookup-based Convolutional Neural Network
- Collaborative Layer-wise Discriminative Learning in Deep Neural Networks
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Towards Large Scale Geostatistical Methane Monitoring with Part-based Object Detection
- Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- Hierarchical Cross-modal Prompt Learning for Vision-Language Models
- Deep Mixture of Experts via Shallow Embedding
- Visualizing and Understanding Deep Texture Representations
- A New Convolutional Network-in-Network Structure and Its Applications in Skin Detection, Semantic Segmentation, and Artifact Reduction
- Feature Agglomeration Networks for Single Stage Face Detection
- Tracking Randomly Moving Objects on Edge Box Proposals
- Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection
- FollowMe: Efficient Online Min-Cost Flow Tracking with Bounded Memory\n and Computation
- Instance Scale Normalization for image understanding
- Doubly Attentive Transformer Machine Translation
- From Captions to Visual Concepts and Back
- Temporal Action Detection with Structured Segment Networks
- Few-Shot Learning in Video and 3D Object Detection: A Survey
- SOD-YOLO: Enhancing YOLO-Based Detection of Small Objects in UAV Imagery
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Inferring 3D Object Pose in RGB-D Images
- Dense Captioning with Joint Inference and Visual Context
- Beyond Local Search: Tracking Objects Everywhere with Instance-Specific Proposals
- Learning non-maximum suppression
- Oriented Boxes for Accurate Instance Segmentation
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- Semantic Image Cropping
- SlumpGuard: An AI-Powered Real-Time System for Automated Concrete Slump Prediction via Video Analysis
- Measuring the Impact of Rotation Equivariance on Aerial Object Detection
- Augmentation Inside the Network
- Self-supervised pre-training with acoustic configurations for replay spoofing detection
- QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
- DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
- Semi-supervised learning and integration of multi-sequence MR-images for carotid vessel wall and plaque segmentation
- Attribute-Graph: A Graph based approach to Image Ranking
- An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- Learning to detect and localize many objects from few examples
- Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol
- Examining the Impact of Blur on Recognition by Convolutional Networks
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- Soft Proposal Networks for Weakly Supervised Object Localization
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- Unsupervised data augmentation for object detection
- Cross-domain Image Retrieval with a Dual Attribute-aware Ranking Network
- Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
- Automatic Labelling for Low-Light Pedestrian Detection
- A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
- Empowering Manufacturers with Privacy-Preserving AI Tools: A Case Study in Privacy-Preserving Machine Learning to Solve Real-World Problems
- Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
- Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation
- DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
- VisionGuard: Synergistic Framework for Helmet Violation Detection
- Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- Benchmarking KAZE and MCM for Multiclass Classification
- Deformable Part Models are Convolutional Neural Networks
- Impression Network for Video Object Detection
- Cross-Attention Message-Passing Transformers for Code-Agnostic Decoding in 6G Networks
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- CSDN: A Context-Gated Self-Adaptive Detection Network for Real-Time Object Detection
- Sketch2code: Generating a website from a paper mockup
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- AI-driven visual monitoring of industrial assembly tasks
- ROAM: a Rich Object Appearance Model with Application to Rotoscoping
- Heart rate and respiratory rate prediction from noisy real-world smartphone based on Deep Learning methods
- ESRPCB: an Edge guided Super-Resolution model and Ensemble learning for tiny Printed Circuit Board Defect detection
- Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
- GeoRecon: Graph-Level Representation Learning for 3D Molecules via Reconstruction-Based Pretraining
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- FindMeIfYouCan: Bringing Open Set metrics to near , far and farther Out-of-Distribution Object Detection
Related