Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
2015/06/04 by Shaoqing Ren, Kaiming He, Ren, Shaoqing +5 · 1664 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1506.01497
openalex publication_date 2015/06/04 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/30
Abstract
State-of-the-art object detection networks depend on region proposal\nalgorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN\nhave reduced the running time of these detection networks, exposing region\nproposal computation as a bottleneck. In this work, we introduce a Region\nProposal Network (RPN) that shares full-image convolutional features with the\ndetection network, thus enabling nearly cost-free region proposals. An RPN is a\nfully convolutional network that simultaneously predicts object bounds and\nobjectness scores at each position. The RPN is trained end-to-end to generate\nhigh-quality region proposals, which are used by Fast R-CNN for detection. We\nfurther merge RPN and Fast R-CNN into a single network by sharing their\nconvolutional features---using the recently popular terminology of neural\nnetworks with 'attention' mechanisms, the RPN component tells the unified\nnetwork where to look. For the very deep VGG-16 model, our detection system has\na frame rate of 5fps (including all steps) on a GPU, while achieving\nstate-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS\nCOCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015\ncompetitions, Faster R-CNN and RPN are the foundations of the 1st-place winning\nentries in several tracks. Code has been made publicly available.\n
Citations
Cited by
- FcaNet: Frequency Channel Attention Networks
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- Exploring Syn-to-Real Domain Adaptation for Military Target Detection
- CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- DeFloMat: Detection with Flow Matching for Stable and Efficient Generative Object Localization
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Towards High-Level Semantic Intelligence
- Neuromorphic Object Detection: An In-Depth Study and Future Directions
- Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks
- Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Teaching People LLM's Errors and Getting it Right
- Fast SAM2 with Text-Driven Token Pruning
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- Hierarchical Modeling Approach to Fast and Accurate Table Recognition
- Granular-ball Guided Masking: Structure-aware Data Augmentation
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Real-World Adversarial Attacks on RF-Based Drone Detectors
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
- Progressive Learned Image Compression for Machine Perception
- PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
- Event Extraction in Large Language Model
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- YolovN-CBi: A Lightweight and Efficient Architecture for Real-Time Detection of Small UAVs
- Building UI/UX Dataset for Dark Pattern Detection and YOLOv12x-based Real-Time Object Recognition Detection System
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- A two-stream network with global-local feature fusion for bone age assessment
- StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
- FlowDet: Unifying Object Detection and Generative Transport Flows
- SARMAE: Masked Autoencoder for SAR Representation Learning
- LAPX: Lightweight Hourglass Network with Global Context
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- A Comprehensive Safety Metric to Evaluate Perception in Autonomous Systems
- Mimicking Human Visual Development for Learning Robust Image Representations
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection
- CarlaNCAP: A Framework for Quantifying the Safety of Vulnerable Road Users in Infrastructure-Assisted Collective Perception Using EuroNCAP Scenarios
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- Reliable Detection of Minute Targets in High-Resolution Aerial Imagery across Temporal Shifts
- Emerging Standards for Machine-to-Machine Video Coding
- Feature Coding for Scalable Machine Vision
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Microscopic Vehicle Trajectories from Heterogeneous and Area-Based Traffic
- A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
- ROI-Packing: Efficient Region-Based Compression for Machine Vision
- Efficient Feature Compression for Machines with Global Statistics Preservation
- Enabling Next-Generation Consumer Experience with Feature Coding for Machines
- New VVC profiles targeting Feature Coding for Machines
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- Omni-Referring Image Segmentation
- JOCA: Task-Driven Joint Optimisation of Camera Hardware and Adaptive Camera Control Algorithms
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- Fast SceneScript: Accurate and Efficient Structured Language Model via Multi-Token Prediction
- YOLO and SGBM Integration for Autonomous Tree Branch Detection and Depth Estimation in Radiata Pine Pruning Applications
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- "All You Need" is Not All You Need for a Paper Title: On the Origins of a Scientific Meme
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- When and how to automate image analysis for wildlife monitoring? Guidelines and lessons from a worked example of seabirds in a dynamic coastal environment
- ALDI-ray: Adapting the ALDI Framework for Security X-ray Object Detection
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- BlinkBud: Detecting Hazards from Behind via Sampled Monocular 3D Detection on a Single Earbud
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- Generalized Medical Phrase Grounding
- PAGen: Phase-guided Amplitude Generation for Domain-adaptive Object Detection
- CogEvo-Edu: Cognitive Evolution Educational Multi-Agent Collaborative System
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- Hierarchical Feature Integration for Multi-Signal Automatic Modulation Recognition
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher’s Bias
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Uni-Hema: Unified Model for Digital Hematopathology
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
- ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- Exploring State-of-the-art models for Early Detection of Forest Fires
- Video Object Recognition in Mobile Edge Networks: Local Tracking or Edge Detection?
- Intelligent Image Search Algorithms Fusing Visual Large Models
- HybriDLA: Hybrid Generation for Document Layout Analysis
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
- Robust Physical Adversarial Patches Using Dynamically Optimized Clusters
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
- VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
- State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
- REXO: Indoor Multi-View Radar Object Detection via 3D Bounding Box Diffusion
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- Person Recognition in Aerial Surveillance: A Decade Survey
- StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
- T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners
- Fast Post-Hoc Confidence Fusion for 3-Class Open-Set Aerial Object Detection
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- When CNNs Outperform Transformers and Mambas: Revisiting Deep Architectures for Dental Caries Segmentation
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
- MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
- Uncover and Unlearn Nuisances: Agnostic Fully Test-Time Adaptation
- MixAR: Mixture Autoregressive Image Generation
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- LLM-YOLOMS: Large Language Model-based Semantic Interpretation and Fault Diagnosis for Wind Turbine Components
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
- How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
- Generalizable Blood Cell Detection via Unified Dataset and Faster R-CNN
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Pixel-level Quality Assessment for Oriented Object Detection
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Visual Bridge: Universal Visual Perception Representations Generating
- Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
- SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection
- A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
- SnowyLane: Robust Lane Detection on Snow-covered Rural Roads Using Infrastructural Elements
- Automatic detection of CMEs using synthetically-trained Mask R-CNN
- Individual identification of brown bears using pose-aware metric learning
- MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- CGF-DETR: Cross-Gated Fusion DETR for Enhanced Pneumonia Detection in Chest X-rays
- A Hybrid YOLOv5-SSD IoT-Based Animal Detection System for Durian Plantation Protection
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- Sewer pipeline condition assessment and defect detection using computer vision
- Temporal Reasoning Graph for Activity Recognition
- An Efficient and Generalizable Transfer Learning Method for Weather Condition Detection on Ground Terminals
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles
- Gaussian Combined Distance: A Generic Metric for Object Detection
- Improving Classification of Occluded Objects through Scene Context
- Detecting Unauthorized Vehicles using Deep Learning for Smart Cities: A Case Study on Bangladesh
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- Adversarial Domain Randomization
- Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
- Learning to Refine Object Segments
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Exemplar-Based Open-Set Panoptic Segmentation Network
- Class-specific Anchoring Proposal for 3D Object Recognition in LIDAR and RGB Images
- Hybrid Task Cascade for Instance Segmentation
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
- Bridging Knowledge Graphs to Generate Scene Graphs
- ACP: Automatic Channel Pruning via Clustering and Swarm Intelligence Optimization for CNN
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- A Single-shot Object Detector with Feature Aggragation and Enhancement
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- DeepIris: Iris Recognition Using A Deep Learning Approach
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- Pose-based Modular Network for Human-Object Interaction Detection
- Probabilistic and Geometric Depth: Detecting Objects in Perspective
- BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
- AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss
- Fine-Grained Object Detection over Scientific Document Images with\n Region Embeddings
- Analysing object detectors from the perspective of co-occurring object\n categories
- Temporal Recurrent Networks for Online Action Detection
- Unsupervised Multiple Person Tracking using AutoEncoder-Based Lifted Multicuts
- Crossover Learning for Fast Online Video Instance Segmentation
- Transfer Learning-based Road Damage Detection for Multiple Countries
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Vision-based Robot Manipulation Learning via Human Demonstrations
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Learning Channel Inter-dependencies at Multiple Scales on Dense Networks for Face Recognition
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information
- Classification based Grasp Detection using Spatial Transformer Network
- On the Importance of Visual Context for Data Augmentation in Scene\n Understanding
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Image Captioning: Transforming Objects into Words
- On the Origin of Deep Learning
- MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual\n Network
- Detection as Regression: Certified Object Detection by Median Smoothing
- EfficientDet: Scalable and Efficient Object Detection
- Siamese Box Adaptive Network for Visual Tracking
- A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation
- Fully Convolutional Networks for Panoptic Segmentation
- Are you doing what I say? On modalities alignment in ALFRED
- Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments
- A Solution to Product detection in Densely Packed Scenes
- Data Extraction from Charts via Single Deep Neural Network
- Refinements in Motion and Appearance for Online Multi-Object Tracking
- End-to-End Incremental Learning
- Dynamic Adversarial Patch for Evading Object Detection Models
- Generative Partition Networks for Multi-Person Pose Estimation
- Dynamic Resolution Network
- KL-Divergence-Based Region Proposal Network for Object Detection
- Toward Intelligent Sensing: Intermediate Deep Feature Compression
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- Feedback Attention for Cell Image Segmentation
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Decision-based Universal Adversarial Attack
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- Semantic Foggy Scene Understanding with Synthetic Data
- Improving Object Detection from Scratch via Gated Feature Reuse
- Visual Semantic Reasoning for Image-Text Matching
- SAIA: Split Artificial Intelligence Architecture for Mobile Healthcare System
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- What Vision-Language Models `See' when they See Scenes
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
- Improving Network Slimming with Nonconvex Regularization
- Localization in the Crowd with Topological Constraints
- DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm Retrieval
- An End-to-End Network for Panoptic Segmentation
- A Self Validation Network for Object-Level Human Attention Estimation
- Progressive Coordinate Transforms for Monocular 3D Object Detection
- GhostNet: More Features From Cheap Operations
- TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- The Devil is in the Boundary: Exploiting Boundary Representation for Basis-based Instance Segmentation
- HAMBox: Delving into Online High-quality Anchors Mining for Detecting Outer Faces
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for\n Road Construction Safety Applications
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- Mask Transfiner for High-Quality Instance Segmentation
- Diagnosing Error in Temporal Action Detectors
- Unsupervised Domain Adaptation for Multispectral Pedestrian Detection
- 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving
- Distilling Image Classifiers in Object Detectors
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Involution: Inverting the Inherence of Convolution for Visual Recognition
- AQD: Towards Accurate Fully-Quantized Object Detection
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- Membership Inference Attacks Against Object Detection Models
- Test-Time Adaptive Object Detection with Foundation Model
- Adversarially Robust Quantum Transfer Learning
- Crop yield prediction using machine learning: A systematic literature review
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Siamese Keypoint Prediction Network for Visual Object Tracking
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- TRIE: End-to-End Text Reading and Information Extraction for Document Understanding
- Feature Fusion for Online Mutual Knowledge Distillation
- Delving into the Imbalance of Positive Proposals in Two-stage Object Detection
- Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Deep GrabCut for Object Selection
- Synergistic Neural Forecasting of Air Pollution with Stochastic Sampling
- Building Damage Detection in Satellite Imagery Using Convolutional Neural Networks
- Gastroscopic Panoramic View: Application to Automatic Polyps Detection under Gastroscopy
- A Few-Shot Sequential Approach for Object Counting
- OOS-DSD: Improving Out-of-stock Detection in Retail Images using Auxiliary Tasks
- IoU-aware Single-stage Object Detector for Accurate Localization
- KongNet: A Multi-headed Deep Learning Model for Detection and Classification of Nuclei in Histopathology Images
- Track to Detect and Segment: An Online Multi-Object Tracker
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis Network
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- Weakly Supervised Action Selection Learning in Video
- Resolving the cybersecurity Data Sharing Paradox to scale up\n cybersecurity via a co-production approach towards data sharing
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- Learning Single/Multi-Attribute of Object with Symmetry and Group
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Measuring and Predicting Tag Importance for Image Retrieval
- A Critical Study on Tea Leaf Disease Detection using Deep Learning Techniques
- Doubly Smoothed Density Estimation with Application on Miners' Unsafe Act Detection
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- TensorMask: A Foundation for Dense Object Segmentation
- Syn2Real: A New Benchmark forSynthetic-to-Real Visual Domain Adaptation
- Object Recognition with and without Objects
- DiffusionLane: Diffusion Model for Lane Detection
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- Group-Attention Single-Shot Detector (GA-SSD): Finding Pulmonary Nodules in Large-Scale CT Images
- On Thin Ice: Towards Explainable Conservation Monitoring via Attribution and Perturbations
- Widening and Squeezing: Towards Accurate and Efficient QNNs
- FrameShield: Adversarially Robust Video Anomaly Detection
- GRAP-MOT: Unsupervised Graph-based Position Weighted Person Multi-camera Multi-object Tracking in a Highly Congested Space
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- FVNet: 3D Front-View Proposal Generation for Real-Time Object Detection from Point Clouds
- Joint Multimedia Event Extraction from Video and Article
- WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Stand-Alone Self-Attention in Vision Models
- Pose Augmentation: Class-agnostic Object Pose Transformation for Object Recognition
- Synthetic Data for Robust Runway Detection
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Physics-Guided Fusion for Robust 3D Tracking of Fast Moving Small Objects
- Real-Time Currency Detection and Voice Feedback for Visually Impaired Individuals
- General Instance Distillation for Object Detection
- A Study of Face Obfuscation in ImageNet
- Detecting Text in Natural Image with Connectionist Text Proposal Network
- Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Automated Morphological Analysis of Neurons in Fluorescence Microscopy Using YOLOv8
- From See to Shield: ML-Assisted Fine-Grained Access Control for Visual Data
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- Temporal Action Localization with Variance-Aware Networks
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- Transferable Active Grasping and Real Embodied Dataset
- Multiple Object Tracking with Mixture Density Networks for Trajectory Estimation
- Cops-Ref: A new Dataset and Task on Compositional Referring Expression Comprehension
- Comparative Analysis of Object Detection Algorithms for Surface Defect Detection
- Learning Human-Object Interaction as Groups
- Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network
- Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis
- Modeling and Analysis of Energy Harvesting and Smart Grid-Powered Wireless Communication Networks: A Contemporary Survey
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning
- Integrating Trustworthy Artificial Intelligence with Energy-Efficient Robotic Arms for Waste Sorting
- A Single Set of Adversarial Clothes Breaks Multiple Defense Methods in the Physical World
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Open-set 3D Object Detection
- Moment-Based Domain Adaptation: Learning Bounds and Algorithms
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- BoardVision: Deployment-ready and Robust Motherboard Defect Detection with YOLO+Faster-RCNN Ensemble
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- A Modular Object Detection System for Humanoid Robots Using YOLO
- Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition Cues
- Universally Slimmable Networks and Improved Training Techniques
- Deep learning-based prediction of response to HER2-targeted neoadjuvant chemotherapy from pre-treatment dynamic breast MRI: A multi-institutional validation study
- Unbiased Scene Graph Generation from Biased Training
- Progressive Cluster Purification for Unsupervised Feature Learning
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- Learning Independent Instance Maps for Crowd Localization
- DocVQA: A Dataset for VQA on Document Images
- Detect Anything via Next Point Prediction
- Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering
- Quantifying the presence of graffiti in urban environments
- A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
- The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data
- Controllable Collision Scenario Generation via Collision Pattern Prediction
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- Post-surgical Endometriosis Segmentation in Laparoscopic Videos
- Visual Question Answering Using Semantic Information from Image\n Descriptions
- NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- Source-Free Object Detection with Detection Transformer
- Implicit Feature Pyramid Network for Object Detection
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis
- The Importance and the Limitations of Sim2Real for Robotic Manipulation in Precision Agriculture
- MRI Brain Tumor Detection with Computer Vision
- PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS with Relationship Recovery
- Learning to Reweight with Deep Interactions
- Towards General Purpose Vision Systems
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- CDSA: Cross-Dimensional Self-Attention for Multivariate, Geo-tagged Time Series Imputation
- SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance Recognition
- Pose Neural Fabrics Search
- Re-ID Driven Localization Refinement for Person Search
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Learning Transferable Adversarial Examples via Ghost Networks
- TARO: Toward Semantically Rich Open-World Object Detection
- Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
- Denoised Diffusion for Object-Focused Image Augmentation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- Towards Adversarially Robust Object Detection
- Structured Knowledge Distillation for Dense Prediction
- Exploring Data Aggregation and Transformations to Generalize across Visual Domains
- Learning to Generate Content-Aware Dynamic Detectors
- Deep Learning based Multi-Modal Sensing for Tracking and State Extraction of Small Quadcopters
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
- Fast-Tracker 2.0: Improving Autonomy of Aerial Tracking with Active Vision and Human Location Regression
- Robust 2D/3D Vehicle Parsing in CVIS
- Curriculum Learning with Diversity for Supervised Computer Vision Tasks
- Spatial-Temporal Block and LSTM Network for Pedestrian Trajectories Prediction
- Curriculum Learning with Synthetic Data for Enhanced Pulmonary Nodule Detection in Chest Radiographs
- Multiple interaction learning with question-type prior knowledge for constraining answer search space in visual question answering
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Region Proposal Network with Graph Prior and IoU-Balance Loss for Landmark Detection in 3D Ultrasound
- EOLO: Embedded Object Segmentation only Look Once
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Tracking by Instance Detection: A Meta-Learning Approach
- G-TAD: Sub-Graph Localization for Temporal Action Detection
- A Unified Object Motion and Affinity Model for Online Multi-Object Tracking
- Two-Stream AMTnet for Action Detection
- Explaining raw data complexity to improve satellite onboard processing
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Getting the Numbers Right\unicodex2014Modelling Multi-Class Object Counting in Dense and Varied Scenes
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Neuroplastic Modular Framework: Cross-Domain Image Classification of Garbage and Industrial Surfaces
- Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization
- Comparative Analysis of YOLOv5, Faster R-CNN, SSD, and RetinaNet for Motorbike Detection in Kigali Autonomous Driving Context
- Improving Transferability of Adversarial Examples with Input Diversity
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- WebVision Database: Visual Learning and Understanding from Web Data
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
- Reformulating HOI Detection as Adaptive Set Prediction
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- Quantum-soft QUBO Suppression for Accurate Object Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Detecting Noteheads in Handwritten Scores with ConvNets and Bounding Box\n Regression
- Video Instance Segmentation
- Referring Expression Comprehension for Small Objects
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- IMAGEdit: Let Any Subject Transform
- Fast Object Detection in Compressed Video
- Group-based Distinctive Image Captioning with Memory Attention
- Improving localization-based approaches for breast cancer screening exam classification
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Momentum Contrast for Unsupervised Visual Representation Learning
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Learning Multi-level Deep Representations for Image Emotion Classification
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- TP-MVCC: Tri-plane Multi-view Fusion Model for Silkie Chicken Counting
- A Fast Knowledge Distillation Framework for Visual Recognition
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
- SPLAT: Semantic Pixel-Level Adaptation Transforms for Detection
- 2nd Place Solution to ECCV 2020 VIPriors Object Detection Challenge
- Semantically-Aware Strategies for Stereo-Visual Robotic Obstacle Avoidance
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
- Diff-3DCap: Shape Captioning with Diffusion Models
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Voxel-FPN: multi-scale voxel feature aggregation in 3D object detection from point clouds
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Enhanced Fracture Diagnosis Based on Critical Regional and Scale Aware in YOLO
- Answer-checking in Context: A Multi-modal FullyAttention Network for Visual Question Answering
- Soft Sampling for Robust Object Detection
- Learning Modulated Loss for Rotated Object Detection
- Tracking Objects as Points
- Exploiting long-term temporal dynamics for video captioning
- One-Shot Object Detection without Fine-Tuning
- TSDM: Tracking by SiamRPN++ with a Depth-refiner and a Mask-generator
- Cross-Modality Relevance for Reasoning on Language and Vision
- Visual Relationship Detection using Scene Graphs: A Survey
- Scope Head for Accurate Localization in Object Detection
- COCAS: A Large-Scale Clothes Changing Person Dataset for Re-identification
- End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection
- γ-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Multilingual Vision-Language Models, A Survey
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Enhancing Vehicle Detection under Adverse Weather Conditions with Contrastive Learning
- Incorporating Scene Context and Semantic Labels for Enhanced Group-level Emotion Recognition
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Deep Watershed Detector for Music Object Recognition
- Learning Student Networks via Feature Embedding
- AI-Enabled Crater-Based Navigation for Lunar Mapping
- FCPose: Fully Convolutional Multi-Person Pose Estimation with Dynamic Instance-Aware Convolutions
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Resolving Class Imbalance in Object Detection with Weighted Cross Entropy Losses
- Visually Grounded Continual Learning of Compositional Phrases
- CompressAI-Vision: Open-source software to evaluate compression methods for computer vision tasks
- Self-Critical Reasoning for Robust Visual Question Answering
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- GridMask Data Augmentation
- DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
- BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
- Monitoring spatial sustainable development: Semi-automated analysis of satellite and aerial images for energy transition and sustainability indicators
- MM-FSOD: Meta and metric integrated few-shot object detection
- The 1st Tiny Object Detection Challenge:Methods and Results
- Not Only Look But Observe: Variational Observation Model of Scene-Level 3D Multi-Object Understanding for Probabilistic SLAM
- Structure-Aware Face Clustering on a Large-Scale Graph with \bf107 Nodes
- One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
- Conditional Convolutions for Instance Segmentation
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- Improving Object Detection with Selective Self-supervised Self-training
- Whole-Body Human Pose Estimation in the Wild
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation
- Towards a Generic Diver-Following Algorithm: Balancing Robustness and\n Efficiency in Deep Visual Detection
- Hard Negative Mixing for Contrastive Learning
- Image Captioning based on Deep Learning Methods: A Survey
- Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
- Semantic Understanding of Scenes Through the ADE20K Dataset
- A CNN approach to simultaneously count plants and detect plantation-rows from UAV imagery
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Multi-view Tracking, Re-ID, and Social Network Analysis of a Flock of Visually Similar Birds in an Outdoor Aviary
- FADE: A Task-Agnostic Upsampling Operator for Encoder–Decoder Architectures
- RevealNet: Seeing Behind Objects in RGB-D Scans
- Semi-Autoregressive Transformer for Image Captioning
- Fast Video Shot Transition Localization with Deep Structured Models
- RiO-DETR: DETR for Real-time Oriented Object Detection
- Gesture Recognition for Initiating Human-to-Robot Handovers
- Forest R-CNN: Large-Vocabulary Long-Tailed Object Detection and Instance Segmentation
- Representation Sharing for Fast Object Detector Search and Beyond
- Where are the Blobs: Counting by Localization with Point Supervision
- Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
- Adaptive Offline Quintuplet Loss for Image-Text Matching
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- IoU Attack: Towards Temporally Coherent Black-Box Adversarial Attack for Visual Object Tracking
- Human-centric Spatio-Temporal Video Grounding With Visual Transformers
- Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
- Deep Affinity Net: Instance Segmentation via Affinity
- ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks
- Bringing in the outliers: A sparse subspace clustering approach to learn a dictionary of mouse ultrasonic vocalizations
- RobustTAD: Robust Time Series Anomaly Detection via Decomposition and Convolutional Neural Networks
- Globally-Aware Multiple Instance Classifier for Breast Cancer Screening
- Learning semantic Image attributes using Image recognition and knowledge graph embeddings
- Object Detection in the Context of Mobile Augmented Reality
- Learning to Discriminate Information for Online Action Detection
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- SOLO: A Simple Framework for Instance Segmentation
- Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology
- Self-supervised 6D Object Pose Estimation for Robot Manipulation
- Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection
- SuperOCR: A Conversion from Optical Character Recognition to Image Captioning
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Unsupervised Part Discovery via Feature Alignment
- Multimodal Learning for Hateful Memes Detection
- Automatic Detection of Cardiac Chambers Using an Attention-based YOLOv4 Framework from Four-chamber View of Fetal Echocardiography
- Positional Encoding as Spatial Inductive Bias in GANs
- ParaNet: Deep Regular Representation for 3D Point Clouds
- Box-Level Class-Balanced Sampling for Active Object Detection
- Robust Segmentation of Optic Disc and Cup from Fundus Images Using Deep Neural Networks
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- From Unstable to Playable: Stabilizing Angry Birds Levels via Object Segmentation
- DRG: Dual Relation Graph for Human-Object Interaction Detection
- Region Refinement Network for Salient Object Detection
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- Physical Adversarial Attack on Vehicle Detector in the Carla Simulator
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- A Deep Ordinal Distortion Estimation Approach for Distortion Rectification
- Language-in-the-Loop Culvert Inspection on the Erie Canal
- Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching
- Searching for Accurate Binary Neural Architectures
- Improving RetinaNet for CT Lesion Detection with Dense Masks from Weak RECIST Labels
- Two-Stream Region Convolutional 3D Network for Temporal Activity Detection
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- Fourier Contour Embedding for Arbitrary-Shaped Text Detection
- Domain Adaptive Object Detection for Space Applications with Real-Time Constraints
- Democratizing Production-Scale Distributed Deep Learning
- FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- An Analysis of Kalman Filter based Object Tracking Methods for Fast-Moving Tiny Objects
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks
- Deep Clustering for Unsupervised Learning of Visual Features
- LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object Detection
- Prototypical Contrastive Learning of Unsupervised Representations
- Enhanced Detection of Tiny Objects in Aerial Images
- From Data to Diagnosis: A Large, Comprehensive Bone Marrow Dataset and AI Methods for Childhood Leukemia Prediction
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- Saccadic Vision for Fine-Grained Visual Classification
- Region-Aware Deformable Convolutions
- Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN
- Rethinking "Batch" in BatchNorm
- PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
- Crafting GBD-Net for Object Detection
- Emotion-Aware Speech Generation with Character-Specific Voices for Comics
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Exploring Self-attention for Image Recognition
- Cross-view Relation Networks for Mammogram Mass Detection
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- NoduleNet: Decoupled False Positive Reductionfor Pulmonary Nodule Detection and Segmentation
- Sequential Voting with Relational Box Fields for Active Object Detection
- Unsupervised Object-Level Representation Learning from Scene Images
- Towards Overcoming False Positives in Visual Relationship Detection
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- Modification method for single-stage object detectors that allows to\n exploit the temporal behaviour of a scene to improve detection accuracy
- An Exploratory Study on Abstract Images and Visual Representations Learned from Them
- Data Leakage in Visual Datasets
- CETUS: Causal Event-Driven Temporal Modeling With Unified Variable-Rate Scheduling
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- Structure Aware SLAM using Quadrics and Planes
- MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
- Simultaneously Localize, Segment and Rank the Camouflaged Objects
- A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- Cumulative Consensus Score: Label-Free and Model-Agnostic Evaluation of Object Detectors in Deployment
- BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation
- Learning in an Uncertain World: Representing Ambiguity Through Multiple\n Hypotheses
- TANet: Robust 3D Object Detection from Point Clouds with Triple Attention
- Robust Object Detection under Occlusion with Context-Aware CompositionalNets
- PPDM: Parallel Point Detection and Matching for Real-time Human-Object Interaction Detection
- Toward Filament Segmentation Using Deep Neural Networks
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- Learning Deep Image Priors for Blind Image Denoising
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- Text-Based Person Search with Limited Data
- SceneGen: Generative Contextual Scene Augmentation using Scene Graph Priors
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Addressing Failure Prediction by Learning Model Confidence
- Bootstrap your own latent: A new approach to self-supervised Learning
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Few-Shot Object Detection via Knowledge Transfer
- Sitatapatra: Blocking the Transfer of Adversarial Samples
- Deep Multi-camera People Detection
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection
- VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations
- A Multiplexed Network for End-to-End, Multilingual OCR
- A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
- Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck
- Optical deep learning nano-profilometry
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- 3D Feature Pyramid Attention Module for Robust Visual Speech Recognition
- Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
- Review: deep learning on 3D point clouds
- Group Evidence Matters: Tiling-based Semantic Gating for Dense Object Detection
- Pose-adaptive Hierarchical Attention Network for Facial Expression Recognition
- Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning
- Deep Learning for Generic Object Detection: A Survey
- Rethinking Channel Dimensions for Efficient Model Design
- Keep it Simple: Image Statistics Matching for Domain Adaptation
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models
- Towards Understanding Visual Grounding in Visual Language Models
- Relation-Aware Graph Attention Network for Visual Question Answering
- Locate, Size and Count: Accurately Resolving People in Dense Crowds via Detection
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- A Co-Training Semi-Supervised Framework Using Faster R-CNN and YOLO Networks for Object Detection in Densely Packed Retail Images
- Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
- Person Re-identification: Past, Present and Future
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- RT-DETR++ for UAV Object Detection
- Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection
- Research on Fast Text Recognition Method for Financial Ticket Image
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- Renovating Parsing R-CNN for Accurate Multiple Human Parsing
- Separating Skills and Concepts for Novel Visual Question Answering
- Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
- Symmetry and Group in Attribute-Object Compositions
- CCL: Cross-modal Correlation Learning with Multi-grained Fusion by Hierarchical Network
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image
- Actor-Centric Relation Network
- Exploring Simple Siamese Representation Learning
- HRank: Filter Pruning using High-Rank Feature Map
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
- Impression Space from Deep Template Network
- Improving Multispectral Pedestrian Detection by Addressing Modality Imbalance Problems
- PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy
- Two-Stage Swarm Intelligence Ensemble Deep Transfer Learning (SI-EDTL) for Vehicle Detection Using Unmanned Aerial Vehicles
- Feature Pyramid Transformer
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- Deep Image Retrieval: Learning global representations for image search
- DeepSEED: 3D Squeeze-and-Excitation Encoder-Decoder Convolutional Neural Networks for Pulmonary Nodule Detection
- Towards Generalization and Data Efficient Learning of Deep Robotic Grasping
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- Towards Robust LiDAR-based Perception in Autonomous Driving: General Black-box Adversarial Sensor Attack and Countermeasures
- FCOS: Fully Convolutional One-Stage Object Detection
- End-to-end Deep Learning Methods for Automated Damage Detection in Extreme Events at Various Scales
- Comprehensive Image Captioning via Scene Graph Decomposition
- Detection of trade in products derived from threatened species using machine learning and a smartphone
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- An Analysis of Scale Invariance in Object Detection - SNIP
- Multi-Modal Camera-Based Detection of Vulnerable Road Users
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Comparison Network for One-Shot Conditional Object Detection
- Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
- Pothole Detection and Recognition based on Transfer Learning
- A Recurrent Vision-and-Language BERT for Navigation
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Semantic Segmentation for Compound figures
- A New Hybrid Model of Generative Adversarial Network and You Only Look Once Algorithm for Automatic License-Plate Recognition
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Detection of E-scooter Riders in Naturalistic Scenes
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
- Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to Sentences
- High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
- FCOS: A simple and strong anchor-free object detector
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
- Bayesian Loss for Crowd Count Estimation with Point Supervision
- Towards Open World Detection: A Survey
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
- TriLiteNet: Lightweight Model for Multi-Task Visual Perception
- Integrating Objects into Monocular SLAM: Line Based Category Specific Models
- Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
- Object Detection in Specific Traffic Scenes using YOLOv2
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- Revisiting the Loss Weight Adjustment in Object Detection
- Dari Barisan ke Pakatan: berubahnya dinamiks Pilihan Raya UmumKuala Lumpur 1955-2013
- Defective Convolutional Networks
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- Semantic Segmentation from Limited Training Data
- Center-based 3D Object Detection and Tracking
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather
- Human De-occlusion: Invisible Perception and Recovery for Humans
- SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
- Heatmap Guided Query Transformers for Robust Astrocyte Detection across Immunostains and Resolutions
- The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
- Vision-based deep execution monitoring
- P2 Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning
- Efficient Pipelines for Vision-Based Context Sensing
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- Lightweight Pyramid Networks for Image Deraining
- Automated data extraction of bar chart raster images
- Global Context Aware RCNN for Object Detection
- Delving into Robust Object Detection from Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement Approach
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- When Healthcare Meets Off-the-Shelf WiFi: A Non-Wearable and Low-Costs Approach for In-Home Monitoring
- The Role of Context Selection in Object Detection
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
- FEDEXCHANGE: Bridging the Domain Gap in Federated Object Detection for Free
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- High-Precision Mixed Feature Fusion Network Using Hypergraph Computation for Cervical Abnormal Cell Detection
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- Unifying Training and Inference for Panoptic Segmentation
- Class-Incremental Few-Shot Object Detection
- GridTracer: Automatic Mapping of Power Grids using Deep Learning and Overhead Imagery
- Towards Deep Learning Assisted Autonomous UAVs for Manipulation Tasks in GPS-Denied Environments
- Trait specialization facilitates autonomous selfing ability in a mixed‐mating plant
- Perceptual Generative Adversarial Networks for Small Object Detection
- Learning from Multiple Datasets with Heterogeneous and Partial Labels for Universal Lesion Detection in CT
- RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction
- COVID-19 personal protective equipment detection using real-time deep learning methods
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- Target-Oriented Single Domain Generalization
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- Understanding the Behaviour of Contrastive Loss
- Learning with Rethinking: Recurrently Improving Convolutional Neural Networks through Feedback
- Probabilistic two-stage detection
- Identifying Surgical Instruments in Laparoscopy Using Deep Learning Instance Segmentation
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in\n Object-oriented SLAM
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- An Empirical Study of DNNs Robustification Inefficacy in Protecting\n Visual Recommenders
- KeepAugment: A Simple Information-Preserving Data Augmentation Approach
- Accurate RGB-D Salient Object Detection via Collaborative Learning
- Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
- IG-TRACK: IOU Guided Siamese Networks for visual object tracking
- Semantic-Aware Ship Detection with Vision-Language Integration
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- YOLOX: Exceeding YOLO Series in 2021
- ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
- Simple Online and Realtime Tracking with a Deep Association Metric
- Learning a Disentangled Embedding for Monocular 3D Shape Retrieval and\n Pose Estimation
- Temporal Action Localization using Long Short-Term Dependency
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
- Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Iterative Low-Rank Approximation for CNN Compression
- What makes instance discrimination good for transfer learning?
- Exploring Randomly Wired Neural Networks for Image Recognition
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- The Role of Teacher Calibration in Knowledge Distillation
- Spatiotemporal Contrastive Video Representation Learning
- Accurate Anchor Free Tracking
- Improving Calibration for Long-Tailed Recognition
- Kernel Transformer Networks for Compact Spherical Convolution
- Objects in Semantic Topology
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Exploring Categorical Regularization for Domain Adaptive Object Detection
- ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network
- ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting
- Learning to Generate Synthetic Data via Compositing
- Co-Separating Sounds of Visual Objects
- LDC-Net: A Unified Framework for Localization, Detection and Counting in Dense Crowds
- Robust and Label-Efficient Deep Waste Detection
- Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Object Detection for Comics using Manga109 Annotations
- Why Do We Click: Visual Impression-aware News Recommendation
- Adaptive Visual Navigation Assistant in 3D RPGs
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- A Deep Journey into Super-resolution: A survey
- Deep Learning Inference in Facebook Data Centers: Characterization,\n Performance Optimizations and Hardware Implications
- AQ-PCDSys: An Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration
- Elucidating image-to-set prediction: An analysis of models, losses and\n datasets
- Revisiting the Sibling Head in Object Detector
- Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models
- Learning Object Relation Graph and Tentative Policy for Visual Navigation
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- Learning to Fuse Things and Stuff
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring
- Understanding top-down attention using task-oriented ablation design
- CLOCs: Camera-LiDAR Object Candidates Fusion for 3D Object Detection
- Incremental Object Detection with Prompt-based Methods
- Toward unsupervised, multi-object discovery in large-scale image\n collections
- Towards Precise End-to-end Weakly Supervised Object Detection Network
- CAD-PU: A Curvature-Adaptive Deep Learning Solution for Point Set Upsampling
- Normalized Cut Loss for Weakly-supervised CNN Segmentation
- AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Revisiting Knowledge Distillation for Object Detection
- CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
- OmViD: Omni-supervised active learning for video action detection
- Large Scale Visual Food Recognition
- Self-Aware Adaptive Alignment: Enabling Accurate Perception for Intelligent Transportation Systems
- Bidirectional Regression for Arbitrary-Shaped Text Detection
- FaPN: Feature-aligned Pyramid Network for Dense Image Prediction
- MF-LPR2: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Gaussian Temporal Awareness Networks for Action Localization
- RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
- Domain Adaptation and Image Classification via Deep Conditional Adaptation Network
- Hybrid Attention for Automatic Segmentation of Whole Fetal Head in Prenatal Ultrasound Volumes
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
- G2L-Net: Global to Local Network for Real-time 6D Pose Estimation with Embedding Vector Features
- supervised adptive threshold network for instance segmentation
- vireoJD-MM at Activity Detection in Extended Videos
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- A system of vision sensor based deep neural networks for complex driving scene analysis in support of crash risk assessment and prevention
- Long Short-Term Relation Networks for Video Action Detection
- Deep Contextual Attention for Human-Object Interaction Detection
- Learning Transferable 3D Adversarial Cloaks for Deep Trained Detectors
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- Data Shift of Object Detection in Autonomous Driving
- Loss re-scaling VQA: Revisiting the LanguagePrior Problem from a Class-imbalance View
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- PEdger++: Practical Edge Detection via Assembling Cross Information
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- MOS: A Low Latency and Lightweight Framework for Face Detection, Landmark Localization, and Head Pose Estimation
- Spatial-Temporal Relation Networks for Multi-Object Tracking
- A Real-time Concrete Crack Detection and Segmentation Model Based on YOLOv11
- Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Two-Level Residual Distillation based Triple Network for Incremental Object Detection
- Adma: A Flexible Loss Function for Neural Networks
- Generic Tubelet Proposals for Action Localization
- A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- Graph-Based Social Relation Reasoning
- Axis-level Symmetry Detection with Group-Equivariant Representation
- EventNet: Asynchronous Recursive Event Processing
- LGA-RCNN: Loss-Guided Attention for Object Detection
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- CSNR and JMIM Based Spectral Band Selection for Reducing Metamerism in Urban Driving
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Towards Powerful and Practical Patch Attacks for 2D Object Detection in Autonomous Driving
- SkeySpot: Automating Service Key Detection for Digital Electrical Layout Plans in the Construction Industry
- CTAP: Complementary Temporal Action Proposal Generation
- Improving OCR for Historical Texts of Multiple Languages
- Glo-UMF: A Unified Multi-model Framework for Automated Morphometry of Glomerular Ultrastructural Characterization
- From Surface to Semantics: Semantic Structure Parsing for Table-Centric Document Analysis
- Image Captioning with Visual Object Representations Grounded in the Textual Modality
- BERT-VQA: Visual Question Answering on Plots
- A Structured Model For Action Detection
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- iFAN: Image-Instance Full Alignment Networks for Adaptive Object Detection
- Shift Equivariance in Object Detection
- Progressive Sparse Local Attention for Video object detection
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Sparse Coding Driven Deep Decision Tree Ensembles for Nuclear Segmentation in Digital Pathology Images
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- MixSearch: Searching for Domain Generalized Medical Image Segmentation Architectures
- SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
- How benign is benign overfitting?
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
- Person Identification with Visual Summary for a Safe Access to a Smart Home
- CFTrack: Center-based Radar and Camera Fusion for 3D Multi-Object Tracking
- Dual-stream Network for Visual Recognition
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- Beyond the Camera: Neural Networks in World Coordinates
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
- Video-based Person Re-identification without Bells and Whistles
- Unsupervised Domain Alignment to Mitigate Low Level Dataset Biases
- SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images
- One-Shot Instance Segmentation
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
- Speaker Diarization with Region Proposal Network
- ZSTAD: Zero-Shot Temporal Activity Detection
- Discovering Visual Patterns in Art Collections with Spatially-consistent Feature Learning
- Self-Supervised Representation Learning for Visual Anomaly Detection
- MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Visual Dialogue State Tracking for Question Generation
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- LungSurg: A Generative AI System for Segmentation and Phase Classification in Thoracoscopic Lobectomy
- LU-Net: a multi-task network to improve the robustness of segmentation of left ventriclular structures by deep learning in 2D echocardiography
- WQT and DG-YOLO: towards domain generalization in underwater object detection
- Training few-shot classification via the perspective of minibatch and pretraining
- Data Priming Network for Automatic Check-Out
- CornerNet: Detecting Objects as Paired Keypoints
- Introducing Pose Consistency and Warp-Alignment for Self-Supervised 6D Object Pose Estimation in Color Images
- Neural Mesh Refiner for 6-DoF Pose Estimation
- WDR FACE: The First Database for Studying Face Detection in Wide Dynamic Range
- Understanding Egocentric Hand-Object Interactions from Hand Pose Estimation
- Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
- CountQA: How Well Do MLLMs Count in the Wild?
- Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction
- Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Segmenting the Complex and Irregular in Two-Phase Flows: A Real-World Empirical Study with SAM2
- Contextual Object Detection with Multimodal Large Language Models
- Parts4Feature: Learning 3D Global Features from Generally Semantic Parts\n in Multiple Views
- Self-Error Adjustment: Theory and Practice of Balancing Individual Performance and Diversity in Ensemble Learning
- Disentangled Deep Autoencoding Regularization for Robust Image Classification
- Benchmarking pig detection and tracking under diverse and challenging conditions
- TopKD: Top-scaled Knowledge Distillation
- A Smartphone-based System for Real-time Early Childhood Caries Diagnosis
- What Holds Back Open-Vocabulary Segmentation?
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- Learning Using Privileged Information for Litter Detection
- CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
- Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters
- Retrieve, Read, Rerank: Towards End-to-End Multi-Document Reading Comprehension
- Deep learning framework for crater detection and identification on the Moon and Mars
- Training Deep Neural Networks via Branch-and-Bound
- GPNAS: A Neural Network Architecture Search Framework Based on Graphical Predictor
- DyCAF-Net: Dynamic Class-Aware Fusion Network
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- Where are the Masks: Instance Segmentation with Image-level Supervision
- The Research of the Real-time Detection and Recognition of Targets in Streetscape Videos
- Temporally Identity-Aware SSD with Attentional LSTM
- Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- TrackNet: A Deep Learning Network for Tracking High-speed and Tiny Objects in Sports Applications
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Online Active Proposal Set Generation for Weakly Supervised Object Detection
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- Feature Selective Small Object Detection via Knowledge-based Recurrent Attentive Neural Network
- CoFF: Cooperative Spatial Feature Fusion for 3D Object Detection on Autonomous Vehicles
- Few-shot Object Detection with Self-adaptive Attention Network for Remote Sensing Images
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- BiFNet: Bidirectional Fusion Network for Road Segmentation
- Towards Compact and Robust Deep Neural Networks
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- Benchmarking Deep Learning-Based Object Detection Models on Feature Deficient Astrophotography Imagery Dataset
- Regula Sub-rosa: Latent Backdoor Attacks on Deep Neural Networks
- InspectVLM: Unified in Theory, Unreliable in Practice
- IAUNet: Instance-Aware U-Net
- TAN: Temporal Affine Network for Real-Time Left Ventricle Anatomical Structure Analysis Based on 2D Ultrasound Videos
- Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
- Spatial-Frequency Aware for Object Detection in RAW Image
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- Geometry-Aware Video Object Detection for Static Cameras
- Multi-Scale Time-Frequency Attention for Acoustic Event Detection
- Screencast-Based Analysis of User-Perceived GUI Responsiveness
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Tobler's First Law in GeoAI: A Spatially Explicit Deep Learning Model for Terrain Feature Detection Under Weak Supervision
- AI-Driven Collaborative Satellite Object Detection for Space Sustainability
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Disrupting Semantic and Abstract Features for Better Adversarial Transferability
- Residual-CNDS for Grand Challenge Scene Dataset
- Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
- CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective
- Backdoor Attacks on Deep Learning Face Detection
- Relation Networks for Object Detection
- DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction
- Fruit Quantity and Quality Estimation using a Robotic Vision System
- Dynamic Temporal Pyramid Network: A Closer Look at Multi-Scale Modeling for Activity Detection
- Human Motion Capture Using a Drone
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- Saliency-Guided Attention Network for Image-Sentence Matching
- Local Dense Logit Relations for Enhanced Knowledge Distillation
- SemiNLL: A Framework of Noisy-Label Learning by Semi-Supervised Learning
- Tricks and Plug-ins for Gradient Boosting in Image Classification
- The Garden of Forking Paths: Towards Multi-Future Trajectory Prediction
- Scalable Change Retrieval Using Deep 3D Neural Codes
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- DDD20 End-to-End Event Camera Driving Dataset: Fusing Frames and Events with Deep Learning for Improved Steering Prediction
- Exploiting Diffusion Prior for Task-driven Image Restoration
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- StarNet: Targeted Computation for Object Detection in Point Clouds
- FDNAS: Improving Data Privacy and Model Diversity in AutoML
- Object Recognition Datasets and Challenges: A Review
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Staining and locking computer vision models without retraining
- SARD: A YOLOv8-Based System for Solar Active Region Detection with SDO/HMI Magnetograms
- Transferring Domain-Agnostic Knowledge in Video Question Answering
- Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection
- Entropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
- RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- Transferable Adversarial Examples for Anchor Free Object Detection
- Automated Detection of Antarctic Benthic Organisms in High-Resolution In Situ Imagery to Aid Biodiversity Monitoring
- Efficient Neural Architecture Transformation Searchin Channel-Level for Object Detection
- Tracking Moose using Aerial Object Detection
- Dual Guidance Semi-Supervised Action Detection
- On Explaining Visual Captioning with Hybrid Markov Logic Networks
- A Frank-Wolfe Framework for Efficient and Effective Adversarial Attacks
- Real-time Visual Object Tracking with Natural Language Description
- Context-Aware Single-Shot Detector
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- Ensemble Foreground Management for Unsupervised Object Discovery
- Regularizing Subspace Redundancy of Low-Rank Adaptation
- Instance Segmentation with Point Supervision
- Object 6D Pose Estimation with Non-local Attention
- Self-Supervised Continuous Colormap Recovery from a 2D Scalar Field Visualization without a Legend
- Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search
- Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning
- An Improved YOLOv8 Approach for Small Target Detection of Rice Spikelet Flowering in Field Environments
- Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications
- TDAPNet: Prototype Network with Recurrent Top-Down Attention for Robust Object Classification under Partial Occlusion
- MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation
- Few-Shot Object Detection via Spatial-Channel State Space Model
- One-Shot Domain Adaptation For Face Generation
- AnimalClue: Recognizing Animals by their Traces
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Towards Universal Modal Tracking with Online Dense Temporal Token Learning
- Region-based Cluster Discrimination for Visual Representation Learning
- Seesaw Loss for Long-Tailed Instance Segmentation
- Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection
- Visual Tracking via Dynamic Memory Networks
- A Large Scale Urban Surveillance Video Dataset for Multiple-Object Tracking and Behavior Analysis
- Adaptive Label Smoothing
- Beef Cattle Instance Segmentation Using Fully Convolutional Neural Network
- RADDet: Range-Azimuth-Doppler based Radar Object Detection for Dynamic Road Users
- A2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Incremental Learning for Robot Perception through HRI
- MITOS-RCNN: A Novel Approach to Mitotic Figure Detection in Breast Cancer Histopathology Images using Region Based Convolutional Neural Networks
- RILOD: Near Real-Time Incremental Learning for Object Detection at the Edge
- Two-phase weakly supervised object detection with pseudo ground truth mining
- CentripetalText: An Efficient Text Instance Representation for Scene\n Text Detection
- Towards Holistic Surgical Scene Graph
- AlteregoNets: a way to human augmentation
- Progressive Stage-wise Learning for Unsupervised Feature Representation Enhancement
- 6-DoF Object Pose from Semantic Keypoints
- Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
- OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
- X-Linear Attention Networks for Image Captioning
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- JDATT: A Joint Distillation Framework for Atmospheric Turbulence Mitigation and Target Detection
- Large Language Model Agent for Structural Drawing Generation Using ReAct Prompt Engineering and Retrieval Augmented Generation
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and beyond
- PhysVarMix: Physics-Informed Variational Mixture Model for Multi-Modal Trajectory Prediction
- Towards a Robust Deep Neural Network in Texts: A Survey
- ABCD: Automatic Blood Cell Detection via Attention-Guided Improved YOLOX
- Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences
- Enhancing Diabetic Retinopathy Classification Accuracy through Dual Attention Mechanism in Deep Learning
- Towards Cross-Modal Forgery Detection and Localization on Live Surveillance Videos
- Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
- WiSE-OD: Benchmarking Robustness in Infrared Object Detection
- Target Driven Instance Detection
- Plot and Rework: Modeling Storylines for Visual Storytelling
- Unsupervised Hard Example Mining from Videos for Improved Object Detection
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Multi-adversarial Faster-RCNN for Unrestricted Object Detection
- A COCO-Formatted Instance-Level Dataset for Plasmodium Falciparum Detection in Giemsa-Stained Blood Smears
- TCM-Tongue: A Standardized Tongue Image Dataset with Pathological Annotations for AI-Assisted TCM Diagnosis
- Few-shot Learning with Global Relatedness Decoupled-Distillation
- Real Time Incremental Foveal Texture Mapping for Autonomous Vehicles
- Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
- Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video
- Towards Generalized and Incremental Few-Shot Object Detection
- Fast query-by-example speech search using separable model
- Learning Context-Aware Embedding for Person Search
- BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking
- Location-Aware Box Reasoning for Anchor-Based Single-Shot Object Detection
- Reverse-engineering Bar Charts Using Neural Networks
- Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
- Towards Large Scale Geostatistical Methane Monitoring with Part-based Object Detection
- Language-Conditioned Imitation Learning for Robot Manipulation Tasks
- Detecting Semantic Parts on Partially Occluded Objects
- 3D Object Proposals using Stereo Imagery for Accurate Object Class Detection
- Object-Centric Unsupervised Image Captioning
- DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
- Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- An End-to-End Foreground-Aware Network for Person Re-Identification
- C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation
- InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
- Distilling Object Detectors with Feature Richness
- SOLO: Segmenting Objects by Locations
- Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019
- Localization-aware Channel Pruning for Object Detection
- Focal and Global Knowledge Distillation for Detectors
- Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- PDPGD: Primal-Dual Proximal Gradient Descent Adversarial Attack
- ShuffleDet: Real-Time Vehicle Detection Network in On-board Embedded UAV\n Imagery
- Spatio-temporal Video Re-localization by Warp LSTM
- A Robotic Approach towards Quantifying Epipelagic Bound Plastic Using Deep Visual Models
- LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- RarePlanes: Synthetic Data Takes Flight
- Pathology-Aware Generative Adversarial Networks for Medical Image Augmentation
- Large Scale Font Independent Urdu Text Recognition System
- Low-light Image Enhancement Algorithm Based on Retinex and Generative Adversarial Network
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- Image De-raining Using a Conditional Generative Adversarial Network
- Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification from the Bottom Up
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- Lightweight Real-time Makeup Try-on in Mobile Browsers with Tiny CNN Models for Facial Tracking
- Context-Aware Zero-Shot Recognition
- Domain Adaptor Networks for Hyperspectral Image Recognition
- Merging Tasks for Video Panoptic Segmentation
- Improving Contrastive Learning by Visualizing Feature Transformation
- Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
- Rethinking Drone-Based Search and Rescue with Aerial Person Detection
- Exploring Semantic Relationships for Unpaired Image Captioning
- TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples
- Saliency deep embedding for aurora image search
- When We First Met: Visual-Inertial Person Localization for Co-Robot Rendezvous
- Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation
- Machine Vision for Improved Human-Robot Cooperation in Adverse Underwater Conditions
- Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection
- Channel-wise Motion Features for Efficient Motion Segmentation
- SECS: Efficient Deep Stream Processing via Class Skew Dichotomy
- Pneumothorax Segmentation: Deep Learning Image Segmentation to predict Pneumothorax
- f-Cal: Calibrated aleatoric uncertainty estimation from neural\n networks for robot perception
- Self-EMD: Self-Supervised Object Detection without ImageNet
- 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection
- Point in, Box out: Beyond Counting Persons in Crowds
- Labeled Data Generation with Inexact Supervision
- PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
- Multilayer Collaborative Low-Rank Coding Network for Robust Deep Subspace Discovery
- ElixirNet: Relation-aware Network Architecture Adaptation for Medical Lesion Detection
- S2DNAS:Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
- Open-World Entity Segmentation
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Image Augmentations for GAN Training
- Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection
- Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
- An End-to-End Framework for Unsupervised Pose Estimation of Occluded Pedestrians
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
- Faceness-Net: Face Detection through Deep Facial Part Responses
- Context and Attribute Grounded Dense Captioning
- Issues in Object Detection in Videos using Common Single-Image CNNs
- Unpaired Pose Guided Human Image Generation
- Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
- BoundarySqueeze: Image Segmentation as Boundary Squeezing
- Open Domain Generalization with Domain-Augmented Meta-Learning
- PP-YOLOv2: A Practical Object Detector
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection
- Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation
- LIT: Light-field Inference of Transparency for Refractive Object Localization
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation
- Semi-convolutional Operators for Instance Segmentation
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- Residual Attention based Network for Hand Bone Age Assessment
- PubLayNet: largest dataset ever for document layout analysis
- HiFT: Hierarchical Feature Transformer for Aerial Tracking
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Empirical Upper Bound in Object Detection and More
- Human Following for Wheeled Robot with Monocular Pan-tilt Camera
- Rethinking Natural Adversarial Examples for Classification Models
- Driver Action Prediction Using Deep (Bidirectional) Recurrent Neural Network
- Detecting, Localising and Classifying Polyps from Colonoscopy Videos using Deep Learning
- Natural Language Video Localization with Learnable Moment Proposals
- Visual Concepts and Compositional Voting
- Reusing Attention for One-stage Lane Topology Understanding
- Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
- Phase Space Reconstruction Network for Lane Intrusion Action Recognition
- Analysis of Plant Nutrient Deficiencies Using Multi-Spectral Imaging and Optimized Segmentation Model
- Joint 2D-3D Breast Cancer Classification
- CPARR: Category-based Proposal Analysis for Referring Relationships
- FishDet-M: A Unified Benchmark for Underwater Fish Detection with CLIP-Guided Model Selection
- Shape Robust Text Detection with Progressive Scale Expansion Network
- Epipolar-Guided Deep Object Matching for Scene Change Detection
- Few-Shot Learning in Video and 3D Object Detection: A Survey
- Can 3D Adversarial Logos Cloak Humans?
- X-volution: On the unification of convolution and self-attention
- Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
- TracKlinic: Diagnosis of Challenge Factors in Visual Tracking
- A Hypergradient Approach to Robust Regression without Correspondence
- Using UAV images and deep learning to enhance the mapping of deadwood in boreal forests
- FishNet: A Camera Localizer using Deep Recurrent Networks
- LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
- Temporal RoI Align for Video Object Recognition
- Interact as You Intend: Intention-Driven Human-Object Interaction Detection
- Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition
- RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
- DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
- CFENet: An Accurate and Efficient Single-Shot Object Detector for Autonomous Driving
- SOD-YOLO: Enhancing YOLO-Based Detection of Small Objects in UAV Imagery
- Few-shot Object Detection with Feature Attention Highlight Module in Remote Sensing Images
- Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection
- Improve bone age assessment by learning from anatomical local regions
- Universal Physical Camouflage Attacks on Object Detectors
- Frequency-Dynamic Attention Modulation for Dense Prediction
- Search Spaces for Neural Model Training
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural Networks
- An Attention Module for Convolutional Neural Networks
- Video Foundation Models for Animal Behavior Analysis
- Attention to Lesion: Lesion-Aware Convolutional Neural Network for Retinal Optical Coherence Tomography Image Classification
- HR-RCNN: Hierarchical Relational Reasoning for Object Detection
- COVR: A test-bed for Visually Grounded Compositional Generalization with real images
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- Location-Aware Feature Selection Text Detection Network
- Prediction-Tracking-Segmentation
- ObjectNet Dataset: Reanalysis and Correction
- Multi-person Articulated Tracking with Spatial and Temporal Embeddings
- ClipCap: CLIP Prefix for Image Captioning
- Exploiting Contextual Information with Deep Neural Networks
- Sperm Detection and Tracking in Phase-Contrast Microscopy Image Sequences using Deep Learning and Modified CSR-DCF
- SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation
- OVANet: One-vs-All Network for Universal Domain Adaptation
- Enhancing Object Detection in Adverse Conditions using Thermal Imaging
- MovieNet: A Holistic Dataset for Movie Understanding
- DIODE: Dilatable Incremental Object Detection
- TOG: Targeted Adversarial Objectness Gradient Attacks on Real-time Object Detection Systems
- CenterMask: single shot instance segmentation with point representation
- Learning Orientation-Estimation Convolutional Neural Network for Building Detection in Optical Remote Sensing Image
- Towards End-to-End Text Spotting in Natural Scenes
- Pointly-Supervised Instance Segmentation
- Face Anti-Spoofing by Learning Polarization Cues in a Real-World Scenario
- Adaptive Object Detection Using Adjacency and Zoom Prediction
- A survey of Object Classification and Detection based on 2D/3D data
- Weakly Supervised Dataset Collection for Robust Person Detection
- Convolutional Neural Networks for User Identificationbased on Motion\n Sensors Represented as Image
- Cascade R-CNN: High Quality Object Detection and Instance Segmentation
- A Multimodal Sentiment Dataset for Video Recommendation
- Adaptive Multi-scale Detection of Acoustic Events
- PSC-Net: Learning Part Spatial Co-occurrence for Occluded Pedestrian Detection
- SIFT Meets CNN: A Decade Survey of Instance Retrieval
- DR Loss: Improving Object Detection by Distributional Ranking
- Face Parsing with RoI Tanh-Warping
- Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object Detectors
- IAN: The Individual Aggregation Network for Person Search
- Sequential End-to-end Network for Efficient Person Search
- Hiding Faces in Plain Sight: Disrupting AI Face Synthesis with Adversarial Perturbations
- Universal Adversarial Perturbations: A Survey
- Stabilized Medical Image Attacks
- Multi-Scale Attention Network for Crowd Counting
- Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks
- NADS-Net: A Nimble Architecture for Driver and Seat Belt Detection via Convolutional Neural Networks
- A Training-free, One-shot Detection Framework For Geospatial Objects In Remote Sensing Images
- Self-Supervised Learning for Semi-Supervised Temporal Action Proposal
- EAdam Optimizer: How ε Impact Adam
- Harvesting, Detecting, and Characterizing Liver Lesions from Large-scale Multi-phase CT Data via Deep Dynamic Texture Learning
- Semantic Image Manipulation Using Scene Graphs
- Relevance Attack on Detectors
- FrostNet: Towards Quantization-Aware Network Architecture Search
- Deep neural networks can be improved using human-derived contextual expectations
- Exploring Low-light Object Detection Techniques
- Boosting ship detection in SAR images with complementary pretraining techniques
- A Video Analysis Method on Wanfang Dataset via Deep Neural Network
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Object-aware Feature Aggregation for Video Object Detection
- Making History Matter: History-Advantage Sequence Training for Visual Dialog
- Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
- Universal Adder Neural Networks
- Panoptic Feature Pyramid Networks
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Scene-Graph Augmented Data-Driven Risk Assessment of Autonomous Vehicle Decisions
- Regularizing Neural Networks via Stochastic Branch Layers
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- LyAm: Robust Non-Convex Optimization for Stable Learning in Noisy Environments
- Learning non-maximum suppression
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- CASIA-Face-Africa: A Large-scale African Face Image Database
- Learning a metacognition for object perception
- On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- PanoNet3D: Combining Semantic and Geometric Understanding for LiDARPoint Cloud Detection
- A Holistically-Guided Decoder for Deep Representation Learning with Applications to Semantic Segmentation and Object Detection
- Dense Multiscale Feature Fusion Pyramid Networks for Object Detection in UAV-Captured Images
- Deformable Tube Network for Action Detection in Videos
- Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and\n Benchmark
- Discovering Spatio-Temporal Action Tubes
- GKNet: Graph-based Keypoints Network for Monocular Pose Estimation of Non-cooperative Spacecraft
- A Comprehensive Survey for Real-World Industrial Defect Detection: Challenges, Approaches, and Prospects
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
- Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
- OOWL500: Overcoming Dataset Collection Bias in the Wild
- Layer-wise Customized Weak Segmentation Block and AIoU Loss for Accurate Object Detection
- CNN-based Density Estimation and Crowd Counting: A Survey
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- 1st Place Solution for the UVO Challenge on Image-based Open-World Segmentation 2021
- Deep Co-Training for Semi-Supervised Image Recognition
- Fine-Grained Zero-Shot Object Detection
- Aggregated Residual Transformations for Deep Neural Networks
- ContourRend: A Segmentation Method for Improving Contours by Rendering
- Perception Improvement for Free: Exploring Imperceptible Black-box Adversarial Attacks on Image Classification
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- SlumpGuard: An AI-Powered Real-Time System for Automated Concrete Slump Prediction via Video Analysis
- Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
- 3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
- Measuring the Impact of Rotation Equivariance on Aerial Object Detection
- Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
- Two is a crowd: tracking relations in videos
- Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training
- LLM-Guided Agentic Object Detection for Open-World Understanding
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs
- Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking
- Channel-wise Alignment for Adaptive Object Detection
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
- TapLab: A Fast Framework for Semantic Video Segmentation Tapping into Compressed-Domain Knowledge
- Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
- RoHOI: Robustness Benchmark for Human-Object Interaction Detection
- Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
- Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
- Move to See Better: Self-Improving Embodied Object Detection
- Knowledge-Enriched Visual Storytelling
- Certainty Driven Consistency Loss on Multi-Teacher Networks for Semi-Supervised Learning
- Visual Semantic Description Generation with MLLMs for Image-Text Matching
- Unified People Tracking with Graph Neural Networks
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- Cross-Resolution SAR Target Detection Using Structural Hierarchy Adaptation and Reliable Adjacency Alignment
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Language and Visual Entity Relationship Graph for Agent Navigation
- Neighbor-view Enhanced Model for Vision and Language Navigation
- Oriented Objects as pairs of Middle Lines
- Towards Detection of Sheep Onboard a UAV
- Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
- 3D-ADAM: A Dataset for 3D Anomaly Detection in Additive Manufacturing
- Rainbow Artifacts from Electromagnetic Signal Injection Attacks on Image Sensors
- A Hybrid Multilayer Extreme Learning Machine for Image Classification with an Application to Quadcopters
- 6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
- Bi-Classifier Determinacy Maximization for Unsupervised Domain Adaptation
- Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
- A model-agnostic active learning approach for animal detection from camera traps
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- LOVON: Legged Open-Vocabulary Object Navigator
- What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
- VisioPath: Vision-Language Enhanced Model Predictive Control for Safe Autonomous Navigation in Mixed Traffic
- Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning
- Vision Based Picking System for Automatic Express Package Dispatching
- DREAM: Document Reconstruction via End-to-end Autoregressive Model
- Learning to Reconstruct and Segment 3D Objects
- SIENet: Spatial Information Enhancement Network for 3D Object Detection from Point Cloud
- DeepApple: Deep Learning-based Apple Detection using a Suppression Mask R-CNN
- Self-Knowledge Distillation with Progressive Refinement of Targets
- Batch Normalization with Enhanced Linear Transformation
- Learning Human-Object Interactions by Graph Parsing Neural Networks
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
- YOLO-APD: Enhancing YOLOv8 for Robust Pedestrian Detection on Complex Road Geometries
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- ReFormer: The Relational Transformer for Image Captioning
- Self-Supervised Real-Time Tracking of Military Vehicles in Low-FPS UAV Footage
- Monocular 3D Object Detection and Box Fitting Trained End-to-End Using\n Intersection-over-Union Loss
- HGNet: High-Order Spatial Awareness Hypergraph and Multi-Scale Context Attention Network for Colorectal Polyp Detection
- Visual identification of individual Holstein-Friesian cattle via deep metric learning
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- LabelEnc: A New Intermediate Supervision Method for Object Detection
- Few-Shot Transformation of Common Actions into Time and Space
- DevMuT: Testing Deep Learning Framework via Developer Expertise-Based Mutation
- Automated Object Behavioral Feature Extraction for Potential Risk Analysis based on Video Sensor
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- Early Convolutions Help Transformers See Better
- ASFormer: Transformer for Action Segmentation
- Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
- Overcoming Statistical Shortcuts for Open-ended Visual Counting
- Self-supervised Object Tracking with Cycle-consistent Siamese Networks
- CUAB: Convolutional Uncertainty Attention Block Enhanced the Chest X-ray Image Analysis
- T-SYNTH: A Knowledge-Based Dataset of Synthetic Breast Images
- Unsupervised data augmentation for object detection
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- Universal Bounding Box Regression and Its Applications
- Answer Them All! Toward Universal Visual Question Answering Models
- MRC-DETR: An Adaptive Multi-Residual Coupled Transformer for Bare Board PCB Defect Detection
- R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous Driving
- C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text Spotting
- SOLQ: Segmenting Objects by Learning Queries
- Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
- Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic Data
- Learning Spatio-Appearance Memory Network for High-Performance Visual Tracking
- Local Metrics for Multi-Object Tracking
- Malignancy Prediction and Lesion Identification from Clinical Dermatological Images
- CircleNet: Anchor-free Detection with Circle Representation
- Automatic Labelling for Low-Light Pedestrian Detection
- FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
- MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
- CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
- A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
- The Problem of Fragmented Occlusion in Object Detection
- DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
- Tomato plant disease detection using transfer learning with C-GAN synthetic images
- Robust Real-Time Pedestrian Detection on Embedded Devices
- Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- Dynamic Refinement Network for Oriented and Densely Packed Object Detection
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
- Domain Adaptation without Source Data
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors
- Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
- Affordance Transfer Learning for Human-Object Interaction Detection
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
- SSA-CNN: Semantic Self-Attention CNN for Pedestrian Detection
- NOCTIS: Novel Object Cyclic Threshold based Instance Segmentation
- Active Measurement: Efficient Estimation at Scale
- Robust Component Detection for Flexible Manufacturing: A Deep Learning Approach to Tray-Free Object Recognition under Variable Lighting
- High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
- UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
- De-Simplifying Pseudo Labels to Enhancing Domain Adaptive Object Detection
- Disentangling Label Distribution for Long-tailed Visual Recognition
- m-RevNet: Deep Reversible Neural Networks with Momentum
- Customizable ROI-Based Deep Image Compression
- Mind the Detail: Uncovering Clinically Relevant Image Details in Accelerated MRI with Semantically Diverse Reconstructions
- Amodal Instance Segmentation
- AFAN: Augmented Feature Alignment Network for Cross-Domain Object Detection
- X-LineNet: Detecting Aircraft in Remote Sensing Images by a pair of Intersecting Line Segments
- A Revised Generative Evaluation of Visual Dialogue
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition
- SelvaBox: A high-resolution dataset for tropical tree crown detection
- Restoring Spatially-Heterogeneous Distortions using Mixture of Experts Network
- Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
- PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
- Event-based Tiny Object Detection: A Benchmark Dataset and Baseline
- Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation
- Semi-Supervised Surface Anomaly Detection of Composite Wind Turbine\n Blades From Drone Imagery
- TOOD: Task-aligned One-stage Object Detection
- Data Augmentation Revisited: Rethinking the Distribution Gap between\n Clean and Augmented Data
- Achieving Human Parity on Visual Question Answering
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- Linear Context Transform Block
- MiniVLM: A Smaller and Faster Vision-Language Model
- OGNet: Salient Object Detection with Output-guided Attention Module
- Congestion Analysis of Convolutional Neural Network-Based Pedestrian Counting Methods on Helicopter Footage
- Person-in-WiFi: Fine-grained Person Perception using WiFi
- CAMERAS: Enhanced Resolution And Sanity preserving Class Activation Mapping for image saliency
- DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
- Dont Even Look Once: Synthesizing Features for Zero-Shot Detection
- Attention to the Burstiness in Visual Prompt Tuning!
- Region-Aware CAM: High-Resolution Weakly-Supervised Defect Segmentation via Salient Region Perception
- A survey of human-in-the-loop for machine learning
- VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
- Dual-Level Collaborative Transformer for Image Captioning
- Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation
- A New Window Loss Function for Bone Fracture Detection and Localization in X-ray Images with Point-based Annotation
- Improving Token-based Object Detection with Video
- Attention-disentangled Uniform Orthogonal Feature Space Optimization for Few-shot Object Detection
- Contextual Heterogeneous Graph Network for Human-Object Interaction Detection
- Learning Human-Object Interaction Detection using Interaction Points
- Social Adaptive Module for Weakly-supervised Group Activity Recognition
- Detecting Text in the Wild with Deep Character Embedding Network
- Disaster mapping from satellites: damage detection with crowdsourced\n point labels
- Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
- Object Tracking Using Spatio-Temporal Future Prediction
- Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Energy Drain of the Object Detection Processing Pipeline for Mobile Devices: Analysis and Implications
- TITAN: Query-Token based Domain Adaptive Adversarial Learning
- VideoMix: Rethinking Data Augmentation for Video Classification
- YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
- Finding Task-Relevant Features for Few-Shot Learning by Category Traversal
- Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology
- 3D for Free: Crossmodal Transfer Learning using HD Maps
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter
- LogoDet-3K: A Large-Scale Image Dataset for Logo Detection
- Class-Agnostic Region-of-Interest Matching in Document Images
- Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- LiDAR point-cloud processing based on projection methods: a comparison
- Implicit Saliency in Deep Neural Networks
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
- Dense Video Captioning using Graph-based Sentence Summarization
- Lightweight Multi-Frame Integration for Robust YOLO Object Detection in Videos
- AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing
- Feature Hallucination for Self-supervised Action Recognition
- From Codicology to Code: A Comparative Study of Transformer and YOLO-based Detectors for Layout Analysis in Historical Documents
- Scalable Dynamic Origin-Destination Demand Estimation Enhanced by High-Resolution Satellite Imagery Data
- Investigating Attention Mechanism in 3D Point Cloud Object Detection
- Geometry Normalization Networks for Accurate Scene Text Detection
- AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
- Memory Enhanced Global-Local Aggregation for Video Object Detection
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- USIS16K: High-Quality Dataset for Underwater Salient Instance Segmentation
- Cooling-Shrinking Attack: Blinding the Tracker with Imperceptible Noises
- Progressive Modality Cooperation for Multi-Modality Domain Adaptation
- Deep Reasoning with Knowledge Graph for Social Relationship Understanding
- Disentangled Motif-aware Graph Learning for Phrase Grounding
- Joint Deep Learning of Facial Expression Synthesis and Recognition
- InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
- USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
- Mask-GD Segmentation Based Robotic Grasp Detection
- Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
- Spatial-temporal Fusion Convolutional Neural Network for Simulated Driving Behavior Recognition
- Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Pattern-Based Phase-Separation of Tracer and Dispersed Phase Particles in Two-Phase Defocusing Particle Tracking Velocimetry
- On the Robustness of Human-Object Interaction Detection against Distribution Shift
- Feedback Driven Multi Stereo Vision System for Real-Time Event Analysis
- Segmentation Mask Guided End-to-End Person Search
- A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & Segmentation
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- CSDN: A Context-Gated Self-Adaptive Detection Network for Real-Time Object Detection
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- HARRISON: A Benchmark on HAshtag Recommendation for Real-world Images in Social Networks
- RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection
- On Feature Normalization and Data Augmentation
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- Unsupervised Pre-training for Person Re-identification
- CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
- Class Agnostic Instance-level Descriptor for Visual Instance Search
- Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
- Convolutional Character Networks
- RealDriveSim: A Realistic Multi-Modal Multi-Task Synthetic Dataset for Autonomous Driving
- End-to-End Video Instance Segmentation with Transformers
- iShape: A First Step Towards Irregular Shape Instance Segmentation
- You Cannot Easily Catch Me: A Low-Detectable Adversarial Patch for Object Detectors
- Hierarchical Scene Parsing by Weakly Supervised Learning with Image Descriptions
- Turning old models fashion again: Recycling classical CNN networks using the Lattice Transformation
- Real-Time Seamless Single Shot 6D Object Pose Prediction
- Opening up Open-World Tracking
- Translate-to-Recognize Networks for RGB-D Scene Recognition
- PointINS: Point-based Instance Segmentation
- 3D Object Detection on Point Clouds using Local Ground-aware and Adaptive Representation of scenes' surface
- AI-driven visual monitoring of industrial assembly tasks
- Image Captioning Based on a Hierarchical Attention Mechanism and Policy Gradient Optimization
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- SAFCAR: Structured Attention Fusion for Compositional Action Recognition
- Open-World Object Counting in Videos
- Vehicle Detection in Deep Learning
- Improving Visual Question Answering by Referring to Generated Paragraph Captions
- Towards Automatic Construction of Diverse, High-quality Image Dataset
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- Rethinking Pseudo-LiDAR Representation
- Learning to Navigate for Fine-grained Classification
- Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems
- GoG: Relation-aware Graph-over-Graph Network for Visual Dialog
- Malaria Detection and Classificaiton
- Multiple Instance Segmentation in Brachial Plexus Ultrasound Image Using BPMSegNet
- Egocentric Human-Object Interaction Detection: A New Benchmark and Method
- YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Hallucination Improves Few-Shot Object Detection
- Robustness Enhancement of Object Detection in Advanced Driver Assistance Systems (ADAS)
Related