Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
2015/06/04 by Shaoqing Ren, Kaiming He, Ren, Shaoqing +5 · 6,279 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Bottleneck #Computation #Computer Vision and Pattern Recognition (cs.CV) #Computer network #Computer science #Computer vision #Convolutional neural network #Data mining #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Frame (networking) #Frame rate #Information retrieval #Merge (version control) #Multimodal Machine Learning Applications #Object detection #Pascal (unit) #Pattern recognition (psychology) #Pooling #cs.CV
paper · pdf · doi:10.48550/arxiv.1506.01497
published in arXiv (Cornell University) 28, 91-99 (Cornell University) · Extended tech report
openalex publication_date 2015/06/04 · arxiv created 2016/01/06 · arxiv updated 2016/01/07 · openalex created_date 2019/06/27 · openalex updated_date 2026/08/08
Abstract
State-of-the-art object detection networks depend on region proposal\nalgorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN\nhave reduced the running time of these detection networks, exposing region\nproposal computation as a bottleneck. In this work, we introduce a Region\nProposal Network (RPN) that shares full-image convolutional features with the\ndetection network, thus enabling nearly cost-free region proposals. An RPN is a\nfully convolutional network that simultaneously predicts object bounds and\nobjectness scores at each position. The RPN is trained end-to-end to generate\nhigh-quality region proposals, which are used by Fast R-CNN for detection. We\nfurther merge RPN and Fast R-CNN into a single network by sharing their\nconvolutional features---using the recently popular terminology of neural\nnetworks with 'attention' mechanisms, the RPN component tells the unified\nnetwork where to look. For the very deep VGG-16 model, our detection system has\na frame rate of 5fps (including all steps) on a GPU, while achieving\nstate-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS\nCOCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015\ncompetitions, Faster R-CNN and RPN are the foundations of the 1st-place winning\nentries in several tracks. Code has been made publicly available.\n
Citations
Cited by
- FcaNet: Frequency Channel Attention Networks
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- Exploring Syn-to-Real Domain Adaptation for Military Target Detection
- CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- DeFloMat: Detection with Flow Matching for Stable and Efficient Generative Object Localization
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Towards High-Level Semantic Intelligence
- Neuromorphic Object Detection: An In-Depth Study and Future Directions
- Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks
- Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Teaching People LLM's Errors and Getting it Right
- Fast SAM2 with Text-Driven Token Pruning
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- Hierarchical Modeling Approach to Fast and Accurate Table Recognition
- Granular Ball Guided Masking: Structure-aware Data Augmentation
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Real-World Adversarial Attacks on RF-Based Drone Detectors
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
- Progressive Learned Image Compression for Machine Perception
- PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
- Event Extraction in Large Language Model
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- YolovN-CBi: A Lightweight and Efficient Architecture for Real-Time Detection of Small UAVs
- Building UI/UX Dataset for Dark Pattern Detection and YOLOv12x-based Real-Time Object Recognition Detection System
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- A two-stream network with global-local feature fusion for bone age assessment
- StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
- FlowDet: Unifying Object Detection and Generative Transport Flows
- SARMAE: Masked Autoencoder for SAR Representation Learning
- LAPX: Lightweight Hourglass Network with Global Context
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- A Comprehensive Safety Metric to Evaluate Perception in Autonomous Systems
- Mimicking Human Visual Development for Learning Robust Image Representations
- LFFD: A Light and Fast Face Detector for Edge Devices
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection
- CarlaNCAP: A Framework for Quantifying the Safety of Vulnerable Road Users in Infrastructure-Assisted Collective Perception Using EuroNCAP Scenarios
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- Reliable Detection of Minute Targets in High-Resolution Aerial Imagery across Temporal Shifts
- Emerging Standards for Machine-to-Machine Video Coding
- Feature Coding for Scalable Machine Vision
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Microscopic Vehicle Trajectories from Heterogeneous and Area-Based Traffic
- A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
- ROI-Packing: Efficient Region-Based Compression for Machine Vision
- Efficient Feature Compression for Machines with Global Statistics Preservation
- Enabling Next-Generation Consumer Experience with Feature Coding for Machines
- New VVC profiles targeting Feature Coding for Machines
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- Omni-Referring Image Segmentation
- JOCA: Task-Driven Joint Optimisation of Camera Hardware and Adaptive Camera Control Algorithms
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- Fast SceneScript: Accurate and Efficient Structured Language Model via Multi-Token Prediction
- YOLO and SGBM Integration for Autonomous Tree Branch Detection and Depth Estimation in Radiata Pine Pruning Applications
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- "All You Need" is Not All You Need for a Paper Title: On the Origins of a Scientific Meme
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- When and how to automate image analysis for wildlife monitoring? Guidelines and lessons from a worked example of seabirds in a dynamic coastal environment
- ALDI-ray: Adapting the ALDI Framework for Security X-ray Object Detection
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- BlinkBud: Detecting Hazards from Behind via Sampled Monocular 3D Detection on a Single Earbud
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- Generalised Medical Phrase Grounding
- PAGen: Phase-guided Amplitude Generation for Domain-adaptive Object Detection
- CogEvo-Edu: Cognitive Evolution Educational Multi-Agent Collaborative System
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- Hierarchical Feature Integration for Multi-Signal Automatic Modulation Recognition
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher’s Bias
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Uni-Hema: Unified Model for Digital Hematopathology
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
- ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- Exploring State-of-the-art models for Early Detection of Forest Fires
- Video Object Recognition in Mobile Edge Networks: Local Tracking or Edge Detection?
- Intelligent Image Search Algorithms Fusing Visual Large Models
- HybriDLA: Hybrid Generation for Document Layout Analysis
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
- Robust Physical Adversarial Patches Using Dynamically Optimized Clusters
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
- On the development of an AI performance and behavioural measures for teaching and classroom management
- VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
- State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
- REXO: Indoor Multi-View Radar Object Detection via 3D Bounding Box Diffusion
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- Person Recognition in Aerial Surveillance: A Decade Survey
- StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
- T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners
- Fast Post-Hoc Confidence Fusion for 3-Class Open-Set Aerial Object Detection
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- When CNNs Outperform Transformers and Mambas: Revisiting Deep Architectures for Dental Caries Segmentation
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
- MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
- Uncover and Unlearn Nuisances: Agnostic Fully Test-Time Adaptation
- MixAR: Mixture Autoregressive Image Generation
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- LLM-YOLOMS: Large Language Model-based Semantic Interpretation and Fault Diagnosis for Wind Turbine Components
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
- How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
- Generalizable Blood Cell Detection via Unified Dataset and Faster R-CNN
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Pixel-level Quality Assessment for Oriented Object Detection
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Visual Bridge: Universal Visual Perception Representations Generating
- Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
- SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
- Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
- SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection
- A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
- SnowyLane: Robust Lane Detection on Snow-covered Rural Roads Using Infrastructural Elements
- Automatic detection of CMEs using synthetically-trained Mask R-CNN
- Individual identification of brown bears using pose-aware metric learning
- MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- CGF-DETR: Cross-Gated Fusion DETR for Enhanced Pneumonia Detection in Chest X-rays
- A Hybrid YOLOv5-SSD IoT-Based Animal Detection System for Durian Plantation Protection
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- Sewer pipeline condition assessment and defect detection using computer vision
- Temporal Reasoning Graph for Activity Recognition
- An Efficient and Generalizable Transfer Learning Method for Weather Condition Detection on Ground Terminals
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles
- Gaussian Combined Distance: A Generic Metric for Object Detection
- Improving Classification of Occluded Objects through Scene Context
- Detecting Unauthorized Vehicles using Deep Learning for Smart Cities: A Case Study on Bangladesh
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- Adversarial Domain Randomization
- Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
- Learning to Refine Object Segments
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Exemplar-Based Open-Set Panoptic Segmentation Network
- Class-specific Anchoring Proposal for 3D Object Recognition in LIDAR and RGB Images
- Hybrid Task Cascade for Instance Segmentation
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
- Bridging Knowledge Graphs to Generate Scene Graphs
- ACP: Automatic Channel Pruning via Clustering and Swarm Intelligence Optimization for CNN
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- A Single-shot Object Detector with Feature Aggragation and Enhancement
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- DeepIris: Iris Recognition Using A Deep Learning Approach
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- Pose-based Modular Network for Human-Object Interaction Detection
- Probabilistic and Geometric Depth: Detecting Objects in Perspective
- BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
- AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss
- Fine-Grained Object Detection over Scientific Document Images with Region Embeddings
- Analysing object detectors from the perspective of co-occurring object categories
- Temporal Recurrent Networks for Online Action Detection
- Unsupervised Multiple Person Tracking using AutoEncoder-Based Lifted Multicuts
- Crossover Learning for Fast Online Video Instance Segmentation
- Transfer Learning-based Road Damage Detection for Multiple Countries
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Vision-based Robot Manipulation Learning via Human Demonstrations
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Learning Channel Inter-dependencies at Multiple Scales on Dense Networks for Face Recognition
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information
- Classification based Grasp Detection using Spatial Transformer Network
- On the Importance of Visual Context for Data Augmentation in Scene\n Understanding
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Image Captioning: Transforming Objects into Words
- On the Origin of Deep Learning
- MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network
- Detection as Regression: Certified Object Detection by Median Smoothing
- EfficientDet: Scalable and Efficient Object Detection
- Siamese Box Adaptive Network for Visual Tracking
- A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation
- Fully Convolutional Networks for Panoptic Segmentation
- Are you doing what I say? On modalities alignment in ALFRED
- Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments
- A Solution to Product detection in Densely Packed Scenes
- Data Extraction from Charts via Single Deep Neural Network
- Refinements in Motion and Appearance for Online Multi-Object Tracking
- End-to-End Incremental Learning
- Dynamic Adversarial Patch for Evading Object Detection Models
- Generative Partition Networks for Multi-Person Pose Estimation
- Dynamic Resolution Network
- KL-Divergence-Based Region Proposal Network for Object Detection
- Toward Intelligent Sensing: Intermediate Deep Feature Compression
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- Feedback Attention for Cell Image Segmentation
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Decision-based Universal Adversarial Attack
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- Semantic Foggy Scene Understanding with Synthetic Data
- Improving Object Detection from Scratch via Gated Feature Reuse
- Visual Semantic Reasoning for Image-Text Matching
- SAIA: Split Artificial Intelligence Architecture for Mobile Healthcare System
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- What Vision-Language Models `See' when they See Scenes
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
- Improving Network Slimming with Nonconvex Regularization
- Localization in the Crowd with Topological Constraints
- DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm Retrieval
- An End-to-End Network for Panoptic Segmentation
- A Self Validation Network for Object-Level Human Attention Estimation
- Progressive Coordinate Transforms for Monocular 3D Object Detection
- GhostNet: More Features From Cheap Operations
- TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- The Devil is in the Boundary: Exploiting Boundary Representation for Basis-based Instance Segmentation
- HAMBox: Delving into Online High-quality Anchors Mining for Detecting Outer Faces
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for Road Construction Safety Applications
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- Mask Transfiner for High-Quality Instance Segmentation
- Diagnosing Error in Temporal Action Detectors
- Unsupervised Domain Adaptation for Multispectral Pedestrian Detection
- 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving
- Distilling Image Classifiers in Object Detectors
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Involution: Inverting the Inherence of Convolution for Visual Recognition
- AQD: Towards Accurate Fully-Quantized Object Detection
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- Membership Inference Attacks Against Object Detection Models
- Test-Time Adaptive Object Detection with Foundation Model
- Adversarially Robust Quantum Transfer Learning
- Crop yield prediction using machine learning: A systematic literature review
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Siamese Keypoint Prediction Network for Visual Object Tracking
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- TRIE: End-to-End Text Reading and Information Extraction for Document Understanding
- Feature Fusion for Online Mutual Knowledge Distillation
- Delving into the Imbalance of Positive Proposals in Two-stage Object Detection
- Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Deep GrabCut for Object Selection
- Synergistic Neural Forecasting of Air Pollution with Stochastic Sampling
- Building Damage Detection in Satellite Imagery Using Convolutional Neural Networks
- Gastroscopic Panoramic View: Application to Automatic Polyps Detection under Gastroscopy
- A Few-Shot Sequential Approach for Object Counting
- OOS-DSD: Improving Out-of-stock Detection in Retail Images using Auxiliary Tasks
- IoU-aware Single-stage Object Detector for Accurate Localization
- KongNet: A Multi-headed Deep Learning Model for Detection and Classification of Nuclei in Histopathology Images
- Track to Detect and Segment: An Online Multi-Object Tracker
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis Network
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- Weakly Supervised Action Selection Learning in Video
- Resolving the cybersecurity Data Sharing Paradox to scale up cybersecurity via a co-production approach towards data sharing
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- Learning Single/Multi-Attribute of Object with Symmetry and Group
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Measuring and Predicting Tag Importance for Image Retrieval
- A Critical Study on Tea Leaf Disease Detection using Deep Learning Techniques
- Doubly Smoothed Density Estimation with Application on Miners' Unsafe Act Detection
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- TensorMask: A Foundation for Dense Object Segmentation
- Syn2Real: A New Benchmark forSynthetic-to-Real Visual Domain Adaptation
- Object Recognition with and without Objects
- DiffusionLane: Diffusion Model for Lane Detection
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- Group-Attention Single-Shot Detector (GA-SSD): Finding Pulmonary Nodules in Large-Scale CT Images
- On Thin Ice: Towards Explainable Conservation Monitoring via Attribution and Perturbations
- Widening and Squeezing: Towards Accurate and Efficient QNNs
- FrameShield: Adversarially Robust Video Anomaly Detection
- GRAP-MOT: Unsupervised Graph-based Position Weighted Person Multi-camera Multi-object Tracking in a Highly Congested Space
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- FVNet: 3D Front-View Proposal Generation for Real-Time Object Detection from Point Clouds
- Joint Multimedia Event Extraction from Video and Article
- WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Stand-Alone Self-Attention in Vision Models
- Pose Augmentation: Class-agnostic Object Pose Transformation for Object Recognition
- Synthetic Data for Robust Runway Detection
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Physics-Guided Fusion for Robust 3D Tracking of Fast Moving Small Objects
- Real-Time Currency Detection and Voice Feedback for Visually Impaired Individuals
- General Instance Distillation for Object Detection
- A Study of Face Obfuscation in ImageNet
- Detecting Text in Natural Image with Connectionist Text Proposal Network
- Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Automated Morphological Analysis of Neurons in Fluorescence Microscopy Using YOLOv8
- From See to Shield: ML-Assisted Fine-Grained Access Control for Visual Data
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- Temporal Action Localization with Variance-Aware Networks
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- Transferable Active Grasping and Real Embodied Dataset
- Multiple Object Tracking with Mixture Density Networks for Trajectory Estimation
- Cops-Ref: A new Dataset and Task on Compositional Referring Expression Comprehension
- Comparative Analysis of Object Detection Algorithms for Surface Defect Detection
- Learning Human-Object Interaction as Groups
- Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network
- Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
- Modeling and Analysis of Energy Harvesting and Smart Grid-Powered Wireless Communication Networks: A Contemporary Survey
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning
- Integrating Trustworthy Artificial Intelligence with Energy-Efficient Robotic Arms for Waste Sorting
- A Single Set of Adversarial Clothes Breaks Multiple Defense Methods in the Physical World
- MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Open-set 3D Object Detection
- Moment-Based Domain Adaptation: Learning Bounds and Algorithms
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- BoardVision: Deployment-ready and Robust Motherboard Defect Detection with YOLO+Faster-RCNN Ensemble
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- A Modular Object Detection System for Humanoid Robots Using YOLO
- Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition Cues
- Universally Slimmable Networks and Improved Training Techniques
- Deep learning-based prediction of response to HER2-targeted neoadjuvant chemotherapy from pre-treatment dynamic breast MRI: A multi-institutional validation study
- Unbiased Scene Graph Generation from Biased Training
- Progressive Cluster Purification for Unsupervised Feature Learning
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- Learning Independent Instance Maps for Crowd Localization
- DocVQA: A Dataset for VQA on Document Images
- Detect Anything via Next Point Prediction
- Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering
- Quantifying the presence of graffiti in urban environments
- A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
- The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data
- Controllable Collision Scenario Generation via Collision Pattern Prediction
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- Post-surgical Endometriosis Segmentation in Laparoscopic Videos
- Visual Question Answering Using Semantic Information from Image Descriptions
- NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- Source-Free Object Detection with Detection Transformer
- Implicit Feature Pyramid Network for Object Detection
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis
- The Importance and the Limitations of Sim2Real for Robotic Manipulation in Precision Agriculture
- MRI Brain Tumor Detection with Computer Vision
- PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS with Relationship Recovery
- Learning to Reweight with Deep Interactions
- Towards General Purpose Vision Systems
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- CDSA: Cross-Dimensional Self-Attention for Multivariate, Geo-tagged Time Series Imputation
- SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance Recognition
- Pose Neural Fabrics Search
- Re-ID Driven Localization Refinement for Person Search
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Learning Transferable Adversarial Examples via Ghost Networks
- TARO: Toward Semantically Rich Open-World Object Detection
- Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
- Denoised Diffusion for Object-Focused Image Augmentation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- Towards Adversarially Robust Object Detection
- Structured Knowledge Distillation for Dense Prediction
- Exploring Data Aggregation and Transformations to Generalize across Visual Domains
- Learning to Generate Content-Aware Dynamic Detectors
- Deep Learning based Multi-Modal Sensing for Tracking and State Extraction of Small Quadcopters
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
- Fast-Tracker 2.0: Improving Autonomy of Aerial Tracking with Active Vision and Human Location Regression
- Robust 2D/3D Vehicle Parsing in CVIS
- Curriculum Learning with Diversity for Supervised Computer Vision Tasks
- Spatial-Temporal Block and LSTM Network for Pedestrian Trajectories Prediction
- Curriculum Learning with Synthetic Data for Enhanced Pulmonary Nodule Detection in Chest Radiographs
- Multiple interaction learning with question-type prior knowledge for constraining answer search space in visual question answering
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Region Proposal Network with Graph Prior and IoU-Balance Loss for Landmark Detection in 3D Ultrasound
- EOLO: Embedded Object Segmentation only Look Once
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Tracking by Instance Detection: A Meta-Learning Approach
- G-TAD: Sub-Graph Localization for Temporal Action Detection
- A Unified Object Motion and Affinity Model for Online Multi-Object Tracking
- Two-Stream AMTnet for Action Detection
- Explaining raw data complexity to improve satellite onboard processing
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Getting the Numbers Right\unicodex2014Modelling Multi-Class Object Counting in Dense and Varied Scenes
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Neuroplastic Modular Framework: Cross-Domain Image Classification of Garbage and Industrial Surfaces
- Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization
- Comparative Analysis of YOLOv5, Faster R-CNN, SSD, and RetinaNet for Motorbike Detection in Kigali Autonomous Driving Context
- Improving Transferability of Adversarial Examples with Input Diversity
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- WebVision Database: Visual Learning and Understanding from Web Data
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
- Reformulating HOI Detection as Adaptive Set Prediction
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- Quantum-soft QUBO Suppression for Accurate Object Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Detecting Noteheads in Handwritten Scores with ConvNets and Bounding Box Regression
- Video Instance Segmentation
- Referring Expression Comprehension for Small Objects
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- IMAGEdit: Let Any Subject Transform
- Fast Object Detection in Compressed Video
- Group-based Distinctive Image Captioning with Memory Attention
- Improving localization-based approaches for breast cancer screening exam classification
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Momentum Contrast for Unsupervised Visual Representation Learning
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Learning Multi-level Deep Representations for Image Emotion Classification
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- TP-MVCC: Tri-plane Multi-view Fusion Model for Silkie Chicken Counting
- A Fast Knowledge Distillation Framework for Visual Recognition
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
- SPLAT: Semantic Pixel-Level Adaptation Transforms for Detection
- 2nd Place Solution to ECCV 2020 VIPriors Object Detection Challenge
- Semantically-Aware Strategies for Stereo-Visual Robotic Obstacle Avoidance
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
- Diff-3DCap: Shape Captioning with Diffusion Models
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Voxel-FPN: multi-scale voxel feature aggregation in 3D object detection from point clouds
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Enhanced Fracture Diagnosis Based on Critical Regional and Scale Aware in YOLO
- Answer-checking in Context: A Multi-modal FullyAttention Network for Visual Question Answering
- Soft Sampling for Robust Object Detection
- Learning Modulated Loss for Rotated Object Detection
- Tracking Objects as Points
- Exploiting long-term temporal dynamics for video captioning
- One-Shot Object Detection without Fine-Tuning
- TSDM: Tracking by SiamRPN++ with a Depth-refiner and a Mask-generator
- Cross-Modality Relevance for Reasoning on Language and Vision
- Visual Relationship Detection using Scene Graphs: A Survey
- Scope Head for Accurate Localization in Object Detection
- COCAS: A Large-Scale Clothes Changing Person Dataset for Re-identification
- End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection
- γ-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Multilingual Vision-Language Models, A Survey
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Enhancing Vehicle Detection under Adverse Weather Conditions with Contrastive Learning
- Incorporating Scene Context and Semantic Labels for Enhanced Group-level Emotion Recognition
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Deep Watershed Detector for Music Object Recognition
- Learning Student Networks via Feature Embedding
- AI-Enabled Crater-Based Navigation for Lunar Mapping
- FCPose: Fully Convolutional Multi-Person Pose Estimation with Dynamic Instance-Aware Convolutions
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Resolving Class Imbalance in Object Detection with Weighted Cross Entropy Losses
- Visually Grounded Continual Learning of Compositional Phrases
- CompressAI-Vision: Open-source software to evaluate compression methods for computer vision tasks
- Self-Critical Reasoning for Robust Visual Question Answering
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- SDE-DET: A Precision Network for Shatian Pomelo Detection in Complex Orchard Environments
- GridMask Data Augmentation
- DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
- BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
- Monitoring spatial sustainable development: Semi-automated analysis of satellite and aerial images for energy transition and sustainability indicators
- MM-FSOD: Meta and metric integrated few-shot object detection
- The 1st Tiny Object Detection Challenge:Methods and Results
- Not Only Look But Observe: Variational Observation Model of Scene-Level 3D Multi-Object Understanding for Probabilistic SLAM
- Structure-Aware Face Clustering on a Large-Scale Graph with \bf107 Nodes
- One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
- Conditional Convolutions for Instance Segmentation
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- Improving Object Detection with Selective Self-supervised Self-training
- Whole-Body Human Pose Estimation in the Wild
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation
- Towards a Generic Diver-Following Algorithm: Balancing Robustness and Efficiency in Deep Visual Detection
- Hard Negative Mixing for Contrastive Learning
- Image Captioning based on Deep Learning Methods: A Survey
- Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
- Recurrence along Depth: Deep Convolutional Neural Networks with Recurrent Layer Aggregation
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
- Semantic Understanding of Scenes Through the ADE20K Dataset
- A CNN Approach to Simultaneously Count Plants and Detect Plantation-Rows from UAV Imagery
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Multi-view Tracking, Re-ID, and Social Network Analysis of a Flock of Visually Similar Birds in an Outdoor Aviary
- FADE: A Task-Agnostic Upsampling Operator for Encoder–Decoder Architectures
- RevealNet: Seeing Behind Objects in RGB-D Scans
- Semi-Autoregressive Transformer for Image Captioning
- Fast Video Shot Transition Localization with Deep Structured Models
- RiO-DETR: DETR for Real-time Oriented Object Detection
- Gesture Recognition for Initiating Human-to-Robot Handovers
- Forest R-CNN: Large-Vocabulary Long-Tailed Object Detection and Instance Segmentation
- Representation Sharing for Fast Object Detector Search and Beyond
- Where are the Blobs: Counting by Localization with Point Supervision
- Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
- Adaptive Offline Quintuplet Loss for Image-Text Matching
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- IoU Attack: Towards Temporally Coherent Black-Box Adversarial Attack for Visual Object Tracking
- Human-centric Spatio-Temporal Video Grounding With Visual Transformers
- Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
- Deep Affinity Net: Instance Segmentation via Affinity
- ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks
- Bringing in the outliers: A sparse subspace clustering approach to learn a dictionary of mouse ultrasonic vocalizations
- RobustTAD: Robust Time Series Anomaly Detection via Decomposition and Convolutional Neural Networks
- Globally-Aware Multiple Instance Classifier for Breast Cancer Screening
- Learning semantic Image attributes using Image recognition and knowledge graph embeddings
- Object Detection in the Context of Mobile Augmented Reality
- Learning to Discriminate Information for Online Action Detection
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- SOLO: A Simple Framework for Instance Segmentation
- Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology
- Self-supervised 6D Object Pose Estimation for Robot Manipulation
- Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection
- SuperOCR: A Conversion from Optical Character Recognition to Image Captioning
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Unsupervised Part Discovery via Feature Alignment
- Multimodal Learning for Hateful Memes Detection
- Automatic Detection of Cardiac Chambers Using an Attention-based YOLOv4 Framework from Four-chamber View of Fetal Echocardiography
- Positional Encoding as Spatial Inductive Bias in GANs
- ParaNet: Deep Regular Representation for 3D Point Clouds
- Box-Level Class-Balanced Sampling for Active Object Detection
- Robust Segmentation of Optic Disc and Cup from Fundus Images Using Deep Neural Networks
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- From Unstable to Playable: Stabilizing Angry Birds Levels via Object Segmentation
- DRG: Dual Relation Graph for Human-Object Interaction Detection
- Region Refinement Network for Salient Object Detection
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- Physical Adversarial Attack on Vehicle Detector in the Carla Simulator
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- A Deep Ordinal Distortion Estimation Approach for Distortion Rectification
- Language-in-the-Loop Culvert Inspection on the Erie Canal
- Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching
- Searching for Accurate Binary Neural Architectures
- Improving RetinaNet for CT Lesion Detection with Dense Masks from Weak RECIST Labels
- Two-Stream Region Convolutional 3D Network for Temporal Activity Detection
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- Fourier Contour Embedding for Arbitrary-Shaped Text Detection
- Domain Adaptive Object Detection for Space Applications with Real-Time Constraints
- Democratizing Production-Scale Distributed Deep Learning
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- An Analysis of Kalman Filter based Object Tracking Methods for Fast-Moving Tiny Objects
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks
- Deep Clustering for Unsupervised Learning of Visual Features
- LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object Detection
- Prototypical Contrastive Learning of Unsupervised Representations
- Enhanced Detection of Tiny Objects in Aerial Images
- From Data to Diagnosis: A Large, Comprehensive Bone Marrow Dataset and AI Methods for Childhood Leukemia Prediction
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- Saccadic Vision for Fine-Grained Visual Classification
- Region-Aware Deformable Convolutions
- Maize Seedling Detection Dataset (MSDD): A Curated High-Resolution RGB Dataset for Seedling Maize Detection and Benchmarking with YOLOv9, YOLO11, YOLOv12 and Faster-RCNN
- Rethinking "Batch" in BatchNorm
- PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
- Crafting GBD-Net for Object Detection
- Emotion-Aware Speech Generation with Character-Specific Voices for Comics
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Exploring Self-attention for Image Recognition
- Cross-view Relation Networks for Mammogram Mass Detection
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- NoduleNet: Decoupled False Positive Reductionfor Pulmonary Nodule Detection and Segmentation
- Sequential Voting with Relational Box Fields for Active Object Detection
- Unsupervised Object-Level Representation Learning from Scene Images
- Towards Overcoming False Positives in Visual Relationship Detection
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- Modification method for single-stage object detectors that allows to exploit the temporal behaviour of a scene to improve detection accuracy
- An Exploratory Study on Abstract Images and Visual Representations Learned from Them
- Data Leakage in Visual Datasets
- CETUS: Causal Event-Driven Temporal Modeling With Unified Variable-Rate Scheduling
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- Structure Aware SLAM using Quadrics and Planes
- MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
- Simultaneously Localize, Segment and Rank the Camouflaged Objects
- A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- Cumulative Consensus Score: Label-Free and Model-Agnostic Evaluation of Object Detectors in Deployment
- BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation
- Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
- TANet: Robust 3D Object Detection from Point Clouds with Triple Attention
- Robust Object Detection under Occlusion with Context-Aware CompositionalNets
- PPDM: Parallel Point Detection and Matching for Real-time Human-Object Interaction Detection
- Toward Filament Segmentation Using Deep Neural Networks
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- Learning Deep Image Priors for Blind Image Denoising
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- Text-Based Person Search with Limited Data
- SceneGen: Generative Contextual Scene Augmentation using Scene Graph Priors
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Addressing Failure Prediction by Learning Model Confidence
- Bootstrap your own latent: A new approach to self-supervised Learning
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Few-Shot Object Detection via Knowledge Transfer
- Sitatapatra: Blocking the Transfer of Adversarial Samples
- Deep Multi-camera People Detection
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection
- VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations
- A Multiplexed Network for End-to-End, Multilingual OCR
- A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
- Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck
- Optical deep learning nano-profilometry
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- 3D Feature Pyramid Attention Module for Robust Visual Speech Recognition
- Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
- Review: deep learning on 3D point clouds
- Group Evidence Matters: Tiling-based Semantic Gating for Dense Object Detection
- Pose-adaptive Hierarchical Attention Network for Facial Expression Recognition
- Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning
- Deep Learning for Generic Object Detection: A Survey
- Rethinking Channel Dimensions for Efficient Model Design
- Keep it Simple: Image Statistics Matching for Domain Adaptation
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models
- Towards Understanding Visual Grounding in Visual Language Models
- Relation-Aware Graph Attention Network for Visual Question Answering
- Locate, Size and Count: Accurately Resolving People in Dense Crowds via Detection
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- A Co-Training Semi-Supervised Framework Using Faster R-CNN and YOLO Networks for Object Detection in Densely Packed Retail Images
- Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
- Person Re-identification: Past, Present and Future
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- RT-DETR++ for UAV Object Detection
- Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection
- Research on Fast Text Recognition Method for Financial Ticket Image
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- Renovating Parsing R-CNN for Accurate Multiple Human Parsing
- Separating Skills and Concepts for Novel Visual Question Answering
- Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
- Symmetry and Group in Attribute-Object Compositions
- CCL: Cross-modal Correlation Learning with Multi-grained Fusion by Hierarchical Network
- CEM-FBGTinyDet: Context-Enhanced Foreground Balance with Gradient Tuning for tiny Objects
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image
- Actor-Centric Relation Network
- Exploring Simple Siamese Representation Learning
- HRank: Filter Pruning using High-Rank Feature Map
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
- Impression Space from Deep Template Network
- Improving Multispectral Pedestrian Detection by Addressing Modality Imbalance Problems
- PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy
- Two-Stage Swarm Intelligence Ensemble Deep Transfer Learning (SI-EDTL) for Vehicle Detection Using Unmanned Aerial Vehicles
- Feature Pyramid Transformer
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- Deep Image Retrieval: Learning global representations for image search
- DeepSEED: 3D Squeeze-and-Excitation Encoder-Decoder Convolutional Neural Networks for Pulmonary Nodule Detection
- Towards Generalization and Data Efficient Learning of Deep Robotic Grasping
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- Towards Robust LiDAR-based Perception in Autonomous Driving: General Black-box Adversarial Sensor Attack and Countermeasures
- FCOS: Fully Convolutional One-Stage Object Detection
- End-to-end Deep Learning Methods for Automated Damage Detection in Extreme Events at Various Scales
- Comprehensive Image Captioning via Scene Graph Decomposition
- Detection of trade in products derived from threatened species using machine learning and a smartphone
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- An Analysis of Scale Invariance in Object Detection - SNIP
- Multi-Modal Camera-Based Detection of Vulnerable Road Users
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Comparison Network for One-Shot Conditional Object Detection
- Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
- Pothole Detection and Recognition based on Transfer Learning
- A Recurrent Vision-and-Language BERT for Navigation
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Semantic Segmentation for Compound figures
- A New Hybrid Model of Generative Adversarial Network and You Only Look Once Algorithm for Automatic License-Plate Recognition
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Detection of E-scooter Riders in Naturalistic Scenes
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
- Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to Sentences
- High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
- FCOS: A simple and strong anchor-free object detector
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
- Bayesian Loss for Crowd Count Estimation with Point Supervision
- Towards Open World Detection: A Survey
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
- TriLiteNet: Lightweight Model for Multi-Task Visual Perception
- Integrating Objects into Monocular SLAM: Line Based Category Specific Models
- Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
- Object Detection in Specific Traffic Scenes using YOLOv2
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- Revisiting the Loss Weight Adjustment in Object Detection
- Dari Barisan ke Pakatan: berubahnya dinamiks Pilihan Raya UmumKuala Lumpur 1955-2013
- Defective Convolutional Networks
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- Semantic Segmentation from Limited Training Data
- Center-based 3D Object Detection and Tracking
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather
- Human De-occlusion: Invisible Perception and Recovery for Humans
- SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
- Heatmap Guided Query Transformers for Robust Astrocyte Detection across Immunostains and Resolutions
- The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
- Vision-based deep execution monitoring
- P2 Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning
- Efficient Pipelines for Vision-Based Context Sensing
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- Lightweight Pyramid Networks for Image Deraining
- Automated data extraction of bar chart raster images
- Global Context Aware RCNN for Object Detection
- Delving into Robust Object Detection from Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement Approach
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- When Healthcare Meets Off-the-Shelf WiFi: A Non-Wearable and Low-Costs Approach for In-Home Monitoring
- The Role of Context Selection in Object Detection
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
- FEDEXCHANGE: Bridging the Domain Gap in Federated Object Detection for Free
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- High-Precision Mixed Feature Fusion Network Using Hypergraph Computation for Cervical Abnormal Cell Detection
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- Unifying Training and Inference for Panoptic Segmentation
- Class-Incremental Few-Shot Object Detection
- GridTracer: Automatic Mapping of Power Grids using Deep Learning and Overhead Imagery
- Towards Deep Learning Assisted Autonomous UAVs for Manipulation Tasks in GPS-Denied Environments
- Trait specialization facilitates autonomous selfing ability in a mixed‐mating plant
- Perceptual Generative Adversarial Networks for Small Object Detection
- Learning from Multiple Datasets with Heterogeneous and Partial Labels for Universal Lesion Detection in CT
- RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction
- COVID-19 personal protective equipment detection using real-time deep learning methods
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- Target-Oriented Single Domain Generalization
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- Understanding the Behaviour of Contrastive Loss
- Learning with Rethinking: Recurrently Improving Convolutional Neural Networks through Feedback
- Probabilistic two-stage detection
- Identifying Surgical Instruments in Laparoscopy Using Deep Learning Instance Segmentation
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in\n Object-oriented SLAM
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- An Empirical Study of DNNs Robustification Inefficacy in Protecting Visual Recommenders
- KeepAugment: A Simple Information-Preserving Data Augmentation Approach
- Accurate RGB-D Salient Object Detection via Collaborative Learning
- Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
- IG-TRACK: IOU Guided Siamese Networks for visual object tracking
- Semantic-Aware Ship Detection with Vision-Language Integration
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- YOLOX: Exceeding YOLO Series in 2021
- ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
- Simple Online and Realtime Tracking with a Deep Association Metric
- Learning a Disentangled Embedding for Monocular 3D Shape Retrieval and Pose Estimation
- Temporal Action Localization using Long Short-Term Dependency
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
- Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Iterative Low-Rank Approximation for CNN Compression
- What makes instance discrimination good for transfer learning?
- Exploring Randomly Wired Neural Networks for Image Recognition
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- The Role of Teacher Calibration in Knowledge Distillation
- Spatiotemporal Contrastive Video Representation Learning
- Accurate Anchor Free Tracking
- Improving Calibration for Long-Tailed Recognition
- Kernel Transformer Networks for Compact Spherical Convolution
- Objects in Semantic Topology
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Exploring Categorical Regularization for Domain Adaptive Object Detection
- ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network
- ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting
- Learning to Generate Synthetic Data via Compositing
- Co-Separating Sounds of Visual Objects
- LDC-Net: A Unified Framework for Localization, Detection and Counting in Dense Crowds
- Robust and Label-Efficient Deep Waste Detection
- Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Object Detection for Comics using Manga109 Annotations
- Why Do We Click: Visual Impression-aware News Recommendation
- Adaptive Visual Navigation Assistant in 3D RPGs
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- A Deep Journey into Super-resolution: A survey
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- AQ-PCDSys: An Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration
- Elucidating image-to-set prediction: An analysis of models, losses and datasets
- Revisiting the Sibling Head in Object Detector
- Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models
- Learning Object Relation Graph and Tentative Policy for Visual Navigation
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- Learning to Fuse Things and Stuff
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring
- Understanding top-down attention using task-oriented ablation design
- CLOCs: Camera-LiDAR Object Candidates Fusion for 3D Object Detection
- Incremental Object Detection with Prompt-based Methods
- Toward unsupervised, multi-object discovery in large-scale image collections
- Towards Precise End-to-end Weakly Supervised Object Detection Network
- CAD-PU: A Curvature-Adaptive Deep Learning Solution for Point Set Upsampling
- Normalized Cut Loss for Weakly-supervised CNN Segmentation
- AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Revisiting Knowledge Distillation for Object Detection
- CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
- OmViD: Omni-supervised active learning for video action detection
- Large Scale Visual Food Recognition
- Self-Aware Adaptive Alignment: Enabling Accurate Perception for Intelligent Transportation Systems
- Bidirectional Regression for Arbitrary-Shaped Text Detection
- FaPN: Feature-aligned Pyramid Network for Dense Image Prediction
- MF-LPR2: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Gaussian Temporal Awareness Networks for Action Localization
- RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
- Domain Adaptation and Image Classification via Deep Conditional Adaptation Network
- Hybrid Attention for Automatic Segmentation of Whole Fetal Head in Prenatal Ultrasound Volumes
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
- G2L-Net: Global to Local Network for Real-time 6D Pose Estimation with Embedding Vector Features
- supervised adptive threshold network for instance segmentation
- vireoJD-MM at Activity Detection in Extended Videos
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- A system of vision sensor based deep neural networks for complex driving scene analysis in support of crash risk assessment and prevention
- Long Short-Term Relation Networks for Video Action Detection
- Deep Contextual Attention for Human-Object Interaction Detection
- Learning Transferable 3D Adversarial Cloaks for Deep Trained Detectors
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- Data Shift of Object Detection in Autonomous Driving
- Loss re-scaling VQA: Revisiting the LanguagePrior Problem from a Class-imbalance View
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- PEdger++: Practical Edge Detection via Assembling Cross Information
- TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
- MOS: A Low Latency and Lightweight Framework for Face Detection, Landmark Localization, and Head Pose Estimation
- Spatial-Temporal Relation Networks for Multi-Object Tracking
- A Real-time Concrete Crack Detection and Segmentation Model Based on YOLOv11
- Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Two-Level Residual Distillation based Triple Network for Incremental Object Detection
- Adma: A Flexible Loss Function for Neural Networks
- Generic Tubelet Proposals for Action Localization
- A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- Graph-Based Social Relation Reasoning
- Axis-level Symmetry Detection with Group-Equivariant Representation
- EventNet: Asynchronous Recursive Event Processing
- LGA-RCNN: Loss-Guided Attention for Object Detection
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- CSNR and JMIM Based Spectral Band Selection for Reducing Metamerism in Urban Driving
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Towards Powerful and Practical Patch Attacks for 2D Object Detection in Autonomous Driving
- SkeySpot: Automating Service Key Detection for Digital Electrical Layout Plans in the Construction Industry
- CTAP: Complementary Temporal Action Proposal Generation
- Improving OCR for Historical Texts of Multiple Languages
- Glo-UMF: A Unified Multi-model Framework for Automated Morphometry of Glomerular Ultrastructural Characterization
- From Surface to Semantics: Semantic Structure Parsing for Table-Centric Document Analysis
- Image Captioning with Visual Object Representations Grounded in the Textual Modality
- BERT-VQA: Visual Question Answering on Plots
- A Structured Model For Action Detection
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- iFAN: Image-Instance Full Alignment Networks for Adaptive Object Detection
- Shift Equivariance in Object Detection
- Progressive Sparse Local Attention for Video object detection
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Sparse Coding Driven Deep Decision Tree Ensembles for Nuclear Segmentation in Digital Pathology Images
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- MixSearch: Searching for Domain Generalized Medical Image Segmentation Architectures
- SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
- How benign is benign overfitting?
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
- Person Identification with Visual Summary for a Safe Access to a Smart Home
- CFTrack: Center-based Radar and Camera Fusion for 3D Multi-Object Tracking
- Dual-stream Network for Visual Recognition
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- Beyond the Camera: Neural Networks in World Coordinates
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
- Video-based Person Re-identification without Bells and Whistles
- Unsupervised Domain Alignment to Mitigate Low Level Dataset Biases
- SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images
- One-Shot Instance Segmentation
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
- Speaker Diarization with Region Proposal Network
- ZSTAD: Zero-Shot Temporal Activity Detection
- Discovering Visual Patterns in Art Collections with Spatially-consistent Feature Learning
- Self-Supervised Representation Learning for Visual Anomaly Detection
- MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Visual Dialogue State Tracking for Question Generation
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- LungSurg: A Generative AI System for Segmentation and Phase Classification in Thoracoscopic Lobectomy
- LU-Net: a multi-task network to improve the robustness of segmentation of left ventriclular structures by deep learning in 2D echocardiography
- WQT and DG-YOLO: towards domain generalization in underwater object detection
- Training few-shot classification via the perspective of minibatch and pretraining
- Data Priming Network for Automatic Check-Out
- CornerNet: Detecting Objects as Paired Keypoints
- Introducing Pose Consistency and Warp-Alignment for Self-Supervised 6D Object Pose Estimation in Color Images
- Neural Mesh Refiner for 6-DoF Pose Estimation
- WDR FACE: The First Database for Studying Face Detection in Wide Dynamic Range
- Understanding Egocentric Hand-Object Interactions from Hand Pose Estimation
- Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
- CountQA: How Well Do MLLMs Count in the Wild?
- Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction
- Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Segmenting the Complex and Irregular in Two-Phase Flows: A Real-World Empirical Study with SAM2
- Contextual Object Detection with Multimodal Large Language Models
- Parts4Feature: Learning 3D Global Features from Generally Semantic Parts in Multiple Views
- Self-Error Adjustment: Theory and Practice of Balancing Individual Performance and Diversity in Ensemble Learning
- Disentangled Deep Autoencoding Regularization for Robust Image Classification
- Benchmarking pig detection and tracking under diverse and challenging conditions
- TopKD: Top-scaled Knowledge Distillation
- A Smartphone-based System for Real-time Early Childhood Caries Diagnosis
- What Holds Back Open-Vocabulary Segmentation?
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- Learning Using Privileged Information for Litter Detection
- CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
- Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters
- Retrieve, Read, Rerank: Towards End-to-End Multi-Document Reading Comprehension
- Deep learning framework for crater detection and identification on the Moon and Mars
- Training Deep Neural Networks via Branch-and-Bound
- GPNAS: A Neural Network Architecture Search Framework Based on Graphical Predictor
- DyCAF-Net: Dynamic Class-Aware Fusion Network
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- Where are the Masks: Instance Segmentation with Image-level Supervision
- The Research of the Real-time Detection and Recognition of Targets in Streetscape Videos
- Temporally Identity-Aware SSD with Attentional LSTM
- Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- TrackNet: A Deep Learning Network for Tracking High-speed and Tiny Objects in Sports Applications
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Online Active Proposal Set Generation for Weakly Supervised Object Detection
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- Feature Selective Small Object Detection via Knowledge-based Recurrent Attentive Neural Network
- CoFF: Cooperative Spatial Feature Fusion for 3D Object Detection on Autonomous Vehicles
- Few-shot Object Detection with Self-adaptive Attention Network for Remote Sensing Images
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- BiFNet: Bidirectional Fusion Network for Road Segmentation
- Towards Compact and Robust Deep Neural Networks
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- Benchmarking Deep Learning-Based Object Detection Models on Feature Deficient Astrophotography Imagery Dataset
- Regula Sub-rosa: Latent Backdoor Attacks on Deep Neural Networks
- InspectVLM: Unified in Theory, Unreliable in Practice
- IAUNet: Instance-Aware U-Net
- TAN: Temporal Affine Network for Real-Time Left Ventricle Anatomical Structure Analysis Based on 2D Ultrasound Videos
- Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
- Spatial-Frequency Aware for Object Detection in RAW Image
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- Geometry-Aware Video Object Detection for Static Cameras
- Multi-Scale Time-Frequency Attention for Acoustic Event Detection
- Screencast-Based Analysis of User-Perceived GUI Responsiveness
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Tobler's First Law in GeoAI: A Spatially Explicit Deep Learning Model for Terrain Feature Detection Under Weak Supervision
- AI-Driven Collaborative Satellite Object Detection for Space Sustainability
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Disrupting Semantic and Abstract Features for Better Adversarial Transferability
- Residual-CNDS for Grand Challenge Scene Dataset
- Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
- CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective
- Backdoor Attacks on Deep Learning Face Detection
- Relation Networks for Object Detection
- DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction
- Fruit Quantity and Quality Estimation using a Robotic Vision System
- Dynamic Temporal Pyramid Network: A Closer Look at Multi-Scale Modeling for Activity Detection
- Human Motion Capture Using a Drone
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- Saliency-Guided Attention Network for Image-Sentence Matching
- Local Dense Logit Relations for Enhanced Knowledge Distillation
- SemiNLL: A Framework of Noisy-Label Learning by Semi-Supervised Learning
- Tricks and Plug-ins for Gradient Boosting in Image Classification
- The Garden of Forking Paths: Towards Multi-Future Trajectory Prediction
- Scalable Change Retrieval Using Deep 3D Neural Codes
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- DDD20 End-to-End Event Camera Driving Dataset: Fusing Frames and Events with Deep Learning for Improved Steering Prediction
- Exploiting Diffusion Prior for Task-driven Image Restoration
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- StarNet: Targeted Computation for Object Detection in Point Clouds
- FDNAS: Improving Data Privacy and Model Diversity in AutoML
- Object Recognition Datasets and Challenges: A Review
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Staining and locking computer vision models without retraining
- SARD: A YOLOv8-Based System for Solar Active Region Detection with SDO/HMI Magnetograms
- Transferring Domain-Agnostic Knowledge in Video Question Answering
- Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection
- Entropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
- RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- Transferable Adversarial Examples for Anchor Free Object Detection
- Automated Detection of Antarctic Benthic Organisms in High-Resolution In Situ Imagery to Aid Biodiversity Monitoring
- Efficient Neural Architecture Transformation Searchin Channel-Level for Object Detection
- Tracking Moose using Aerial Object Detection
- Dual Guidance Semi-Supervised Action Detection
- On Explaining Visual Captioning with Hybrid Markov Logic Networks
- A Frank-Wolfe Framework for Efficient and Effective Adversarial Attacks
- Real-time Visual Object Tracking with Natural Language Description
- Context-Aware Single-Shot Detector
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- Ensemble Foreground Management for Unsupervised Object Discovery
- Regularizing Subspace Redundancy of Low-Rank Adaptation
- Instance Segmentation with Point Supervision
- Object 6D Pose Estimation with Non-local Attention
- Self-Supervised Continuous Colormap Recovery from a 2D Scalar Field Visualization without a Legend
- Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search
- Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning
- An Improved YOLOv8 Approach for Small Target Detection of Rice Spikelet Flowering in Field Environments
- Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications
- TDAPNet: Prototype Network with Recurrent Top-Down Attention for Robust Object Classification under Partial Occlusion
- MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation
- Few-Shot Object Detection via Spatial-Channel State Space Model
- One-Shot Domain Adaptation For Face Generation
- AnimalClue: Recognizing Animals by their Traces
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Towards Universal Modal Tracking with Online Dense Temporal Token Learning
- Region-based Cluster Discrimination for Visual Representation Learning
- Seesaw Loss for Long-Tailed Instance Segmentation
- Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection
- Visual Tracking via Dynamic Memory Networks
- A Large Scale Urban Surveillance Video Dataset for Multiple-Object Tracking and Behavior Analysis
- Adaptive Label Smoothing
- Beef Cattle Instance Segmentation Using Fully Convolutional Neural Network
- RADDet: Range-Azimuth-Doppler based Radar Object Detection for Dynamic Road Users
- A2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Incremental Learning for Robot Perception through HRI
- MITOS-RCNN: A Novel Approach to Mitotic Figure Detection in Breast Cancer Histopathology Images using Region Based Convolutional Neural Networks
- RILOD: Near Real-Time Incremental Learning for Object Detection at the Edge
- Two-phase weakly supervised object detection with pseudo ground truth mining
- CentripetalText: An Efficient Text Instance Representation for Scene Text Detection
- Towards Holistic Surgical Scene Graph
- AlteregoNets: a way to human augmentation
- Progressive Stage-wise Learning for Unsupervised Feature Representation Enhancement
- 6-DoF Object Pose from Semantic Keypoints
- Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
- OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
- X-Linear Attention Networks for Image Captioning
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- JDATT: A Joint Distillation Framework for Atmospheric Turbulence Mitigation and Target Detection
- Large Language Model Agent for Structural Drawing Generation Using ReAct Prompt Engineering and Retrieval Augmented Generation
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and beyond
- PhysVarMix: Physics-Informed Variational Mixture Model for Multi-Modal Trajectory Prediction
- Towards a Robust Deep Neural Network in Texts: A Survey
- ABCD: Automatic Blood Cell Detection via Attention-Guided Improved YOLOX
- Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences
- Enhancing Diabetic Retinopathy Classification Accuracy through Dual Attention Mechanism in Deep Learning
- Towards Cross-Modal Forgery Detection and Localization on Live Surveillance Videos
- Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
- YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study
- PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
- WiSE-OD: Benchmarking Robustness in Infrared Object Detection
- Target Driven Instance Detection
- Plot and Rework: Modeling Storylines for Visual Storytelling
- Unsupervised Hard Example Mining from Videos for Improved Object Detection
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Multi-adversarial Faster-RCNN for Unrestricted Object Detection
- A COCO-Formatted Instance-Level Dataset for Plasmodium Falciparum Detection in Giemsa-Stained Blood Smears
- TCM-Tongue: A Standardized Tongue Image Dataset with Pathological Annotations for AI-Assisted TCM Diagnosis
- Few-shot Learning with Global Relatedness Decoupled-Distillation
- Real Time Incremental Foveal Texture Mapping for Autonomous Vehicles
- Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
- Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video
- Towards Generalized and Incremental Few-Shot Object Detection
- Fast query-by-example speech search using separable model
- Learning Context-Aware Embedding for Person Search
- BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking
- Location-Aware Box Reasoning for Anchor-Based Single-Shot Object Detection
- Reverse-engineering Bar Charts Using Neural Networks
- Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
- Towards Large Scale Geostatistical Methane Monitoring with Part-based Object Detection
- Language-Conditioned Imitation Learning for Robot Manipulation Tasks
- Detecting Semantic Parts on Partially Occluded Objects
- 3D Object Proposals using Stereo Imagery for Accurate Object Class Detection
- Object-Centric Unsupervised Image Captioning
- DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
- Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- An End-to-End Foreground-Aware Network for Person Re-Identification
- C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation
- InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
- Distilling Object Detectors with Feature Richness
- SOLO: Segmenting Objects by Locations
- Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019
- Localization-aware Channel Pruning for Object Detection
- Focal and Global Knowledge Distillation for Detectors
- Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- PDPGD: Primal-Dual Proximal Gradient Descent Adversarial Attack
- ShuffleDet: Real-Time Vehicle Detection Network in On-board Embedded UAV Imagery
- Spatio-temporal Video Re-localization by Warp LSTM
- A Robotic Approach towards Quantifying Epipelagic Bound Plastic Using Deep Visual Models
- LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- Physics-based Scene-level Reasoning for Object Pose Estimation in Clutter
- RarePlanes: Synthetic Data Takes Flight
- Pathology-Aware Generative Adversarial Networks for Medical Image Augmentation
- Large Scale Font Independent Urdu Text Recognition System
- Low-light Image Enhancement Algorithm Based on Retinex and Generative Adversarial Network
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- Image De-raining Using a Conditional Generative Adversarial Network
- Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification from the Bottom Up
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- Lightweight Real-time Makeup Try-on in Mobile Browsers with Tiny CNN Models for Facial Tracking
- Context-Aware Zero-Shot Recognition
- Domain Adaptor Networks for Hyperspectral Image Recognition
- Merging Tasks for Video Panoptic Segmentation
- Improving Contrastive Learning by Visualizing Feature Transformation
- Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
- Rethinking Drone-Based Search and Rescue with Aerial Person Detection
- Exploring Semantic Relationships for Unpaired Image Captioning
- TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples
- Saliency deep embedding for aurora image search
- When We First Met: Visual-Inertial Person Localization for Co-Robot Rendezvous
- Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation
- Machine Vision for Improved Human-Robot Cooperation in Adverse Underwater Conditions
- Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection
- Channel-wise Motion Features for Efficient Motion Segmentation
- SECS: Efficient Deep Stream Processing via Class Skew Dichotomy
- Pneumothorax Segmentation: Deep Learning Image Segmentation to predict Pneumothorax
- f-Cal: Calibrated aleatoric uncertainty estimation from neural networks for robot perception
- Self-EMD: Self-Supervised Object Detection without ImageNet
- 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection
- Point in, Box out: Beyond Counting Persons in Crowds
- Labeled Data Generation with Inexact Supervision
- PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
- Multilayer Collaborative Low-Rank Coding Network for Robust Deep Subspace Discovery
- ElixirNet: Relation-aware Network Architecture Adaptation for Medical Lesion Detection
- S2DNAS:Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
- Open-World Entity Segmentation
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Image Augmentations for GAN Training
- Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection
- Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
- An End-to-End Framework for Unsupervised Pose Estimation of Occluded Pedestrians
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
- Faceness-Net: Face Detection through Deep Facial Part Responses
- Context and Attribute Grounded Dense Captioning
- Issues in Object Detection in Videos using Common Single-Image CNNs
- Unpaired Pose Guided Human Image Generation
- Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
- BoundarySqueeze: Image Segmentation as Boundary Squeezing
- Open Domain Generalization with Domain-Augmented Meta-Learning
- PP-YOLOv2: A Practical Object Detector
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection
- Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation
- LIT: Light-field Inference of Transparency for Refractive Object Localization
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation
- Semi-convolutional Operators for Instance Segmentation
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- Residual Attention based Network for Hand Bone Age Assessment
- PubLayNet: largest dataset ever for document layout analysis
- HiFT: Hierarchical Feature Transformer for Aerial Tracking
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Empirical Upper Bound in Object Detection and More
- Human Following for Wheeled Robot with Monocular Pan-tilt Camera
- Rethinking Natural Adversarial Examples for Classification Models
- Driver Action Prediction Using Deep (Bidirectional) Recurrent Neural Network
- Detecting, Localising and Classifying Polyps from Colonoscopy Videos using Deep Learning
- Natural Language Video Localization with Learnable Moment Proposals
- Visual Concepts and Compositional Voting
- Reusing Attention for One-stage Lane Topology Understanding
- Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
- Phase Space Reconstruction Network for Lane Intrusion Action Recognition
- Analysis of Plant Nutrient Deficiencies Using Multi-Spectral Imaging and Optimized Segmentation Model
- Joint 2D-3D Breast Cancer Classification
- CPARR: Category-based Proposal Analysis for Referring Relationships
- FishDet-M: A Unified Benchmark for Underwater Fish Detection with CLIP-Guided Model Selection
- Shape Robust Text Detection with Progressive Scale Expansion Network
- Epipolar-Guided Deep Object Matching for Scene Change Detection
- Few-Shot Learning in Video and 3D Object Detection: A Survey
- Can 3D Adversarial Logos Cloak Humans?
- X-volution: On the unification of convolution and self-attention
- Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
- TracKlinic: Diagnosis of Challenge Factors in Visual Tracking
- A Hypergradient Approach to Robust Regression without Correspondence
- Using UAV images and deep learning to enhance the mapping of deadwood in boreal forests
- FishNet: A Camera Localizer using Deep Recurrent Networks
- AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant T2I Adversarial Patches
- LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
- Temporal RoI Align for Video Object Recognition
- Interact as You Intend: Intention-Driven Human-Object Interaction Detection
- Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition
- Light-Weight RetinaNet for Object Detection
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
- DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
- CFENet: An Accurate and Efficient Single-Shot Object Detector for Autonomous Driving
- Learning Context Graph for Person Search
- SOD-YOLO: Enhancing YOLO-Based Detection of Small Objects in UAV Imagery
- Few-shot Object Detection with Feature Attention Highlight Module in Remote Sensing Images
- Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection
- Improve bone age assessment by learning from anatomical local regions
- Universal Physical Camouflage Attacks on Object Detectors
- Frequency-Dynamic Attention Modulation for Dense Prediction
- Search Spaces for Neural Model Training
- 3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
- Bilinear Graph Networks for Visual Question Answering
- Rethinking of Radar's Role: A Camera-Radar Dataset and Systematic Annotator via Coordinate Alignment
- InterpIoU: Rethinking Bounding Box Regression with Interpolation-Based IoU Optimization
- Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural Networks
- An Attention Module for Convolutional Neural Networks
- Video Foundation Models for Animal Behavior Analysis
- Attention to Lesion: Lesion-Aware Convolutional Neural Network for Retinal Optical Coherence Tomography Image Classification
- HR-RCNN: Hierarchical Relational Reasoning for Object Detection
- COVR: A test-bed for Visually Grounded Compositional Generalization with real images
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- Location-Aware Feature Selection Text Detection Network
- Prediction-Tracking-Segmentation
- ObjectNet Dataset: Reanalysis and Correction
- Multi-person Articulated Tracking with Spatial and Temporal Embeddings
- ClipCap: CLIP Prefix for Image Captioning
- Exploiting Contextual Information with Deep Neural Networks
- Sperm Detection and Tracking in Phase-Contrast Microscopy Image Sequences using Deep Learning and Modified CSR-DCF
- SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation
- OVANet: One-vs-All Network for Universal Domain Adaptation
- Enhancing Object Detection in Adverse Conditions using Thermal Imaging
- MovieNet: A Holistic Dataset for Movie Understanding
- DIODE: Dilatable Incremental Object Detection
- TOG: Targeted Adversarial Objectness Gradient Attacks on Real-time Object Detection Systems
- CenterMask: single shot instance segmentation with point representation
- Learning Orientation-Estimation Convolutional Neural Network for Building Detection in Optical Remote Sensing Image
- Towards End-to-End Text Spotting in Natural Scenes
- Pointly-Supervised Instance Segmentation
- Face Anti-Spoofing by Learning Polarization Cues in a Real-World Scenario
- Adaptive Object Detection Using Adjacency and Zoom Prediction
- A survey of Object Classification and Detection based on 2D/3D data
- Weakly Supervised Dataset Collection for Robust Person Detection
- Convolutional Neural Networks for User Identificationbased on Motion Sensors Represented as Image
- Cascade R-CNN: High Quality Object Detection and Instance Segmentation
- A Multimodal Sentiment Dataset for Video Recommendation
- Adaptive Multi-scale Detection of Acoustic Events
- PSC-Net: Learning Part Spatial Co-occurrence for Occluded Pedestrian Detection
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- SIFT Meets CNN: A Decade Survey of Instance Retrieval
- DR Loss: Improving Object Detection by Distributional Ranking
- Face Parsing with RoI Tanh-Warping
- Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object Detectors
- IAN: The Individual Aggregation Network for Person Search
- Sequential End-to-end Network for Efficient Person Search
- Hiding Faces in Plain Sight: Disrupting AI Face Synthesis with Adversarial Perturbations
- Universal Adversarial Perturbations: A Survey
- Stabilized Medical Image Attacks
- Multi-Scale Attention Network for Crowd Counting
- Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks
- NADS-Net: A Nimble Architecture for Driver and Seat Belt Detection via Convolutional Neural Networks
- A Training-free, One-shot Detection Framework For Geospatial Objects In Remote Sensing Images
- Self-Supervised Learning for Semi-Supervised Temporal Action Proposal
- EAdam Optimizer: How ε Impact Adam
- Harvesting, Detecting, and Characterizing Liver Lesions from Large-scale Multi-phase CT Data via Deep Dynamic Texture Learning
- Semantic Image Manipulation Using Scene Graphs
- Relevance Attack on Detectors
- FrostNet: Towards Quantization-Aware Network Architecture Search
- Deep neural networks can be improved using human-derived contextual expectations
- Exploring Low-light Object Detection Techniques
- Boosting ship detection in SAR images with complementary pretraining techniques
- A Video Analysis Method on Wanfang Dataset via Deep Neural Network
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Object-aware Feature Aggregation for Video Object Detection
- Making History Matter: History-Advantage Sequence Training for Visual Dialog
- Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
- Universal Adder Neural Networks
- Panoptic Feature Pyramid Networks
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Scene-Graph Augmented Data-Driven Risk Assessment of Autonomous Vehicle Decisions
- Regularizing Neural Networks via Stochastic Branch Layers
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- LyAm: Robust Non-Convex Optimization for Stable Learning in Noisy Environments
- Learning non-maximum suppression
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- CASIA-Face-Africa: A Large-scale African Face Image Database
- Learning a metacognition for object perception
- On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- PanoNet3D: Combining Semantic and Geometric Understanding for LiDARPoint Cloud Detection
- A Holistically-Guided Decoder for Deep Representation Learning with Applications to Semantic Segmentation and Object Detection
- Dense Multiscale Feature Fusion Pyramid Networks for Object Detection in UAV-Captured Images
- Deformable Tube Network for Action Detection in Videos
- Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and Benchmark
- Discovering Spatio-Temporal Action Tubes
- GKNet: Graph-based Keypoints Network for Monocular Pose Estimation of Non-cooperative Spacecraft
- A Comprehensive Survey for Real-World Industrial Defect Detection: Challenges, Approaches, and Prospects
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
- Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
- OOWL500: Overcoming Dataset Collection Bias in the Wild
- Layer-wise Customized Weak Segmentation Block and AIoU Loss for Accurate Object Detection
- CNN-based Density Estimation and Crowd Counting: A Survey
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- 1st Place Solution for the UVO Challenge on Image-based Open-World Segmentation 2021
- Deep Co-Training for Semi-Supervised Image Recognition
- Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
- Real-Time Text Detection and Recognition
- Fine-Grained Zero-Shot Object Detection
- Aggregated Residual Transformations for Deep Neural Networks
- ContourRend: A Segmentation Method for Improving Contours by Rendering
- Perception Improvement for Free: Exploring Imperceptible Black-box Adversarial Attacks on Image Classification
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- SlumpGuard: An AI-Powered Real-Time System for Automated Concrete Slump Prediction via Video Analysis
- Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
- 3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
- ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations
- Measuring the Impact of Rotation Equivariance on Aerial Object Detection
- Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
- Two is a crowd: tracking relations in videos
- Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training
- LLM-Guided Agentic Object Detection for Open-World Understanding
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs
- Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking
- Channel-wise Alignment for Adaptive Object Detection
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
- TapLab: A Fast Framework for Semantic Video Segmentation Tapping into Compressed-Domain Knowledge
- Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
- RoHOI: Robustness Benchmark for Human-Object Interaction Detection
- Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
- Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
- Move to See Better: Self-Improving Embodied Object Detection
- Knowledge-Enriched Visual Storytelling
- Certainty Driven Consistency Loss on Multi-Teacher Networks for Semi-Supervised Learning
- Visual Semantic Description Generation with MLLMs for Image-Text Matching
- Unified People Tracking with Graph Neural Networks
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- Cross-Resolution SAR Target Detection Using Structural Hierarchy Adaptation and Reliable Adjacency Alignment
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Language and Visual Entity Relationship Graph for Agent Navigation
- Neighbor-view Enhanced Model for Vision and Language Navigation
- Oriented Objects as pairs of Middle Lines
- Towards Detection of Sheep Onboard a UAV
- Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
- Lightweight Object Detection Using Quantized YOLOv4-Tiny for Emergency Response in Aerial Imagery
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
- 3D-ADAM: A Dataset for 3D Anomaly Detection in Additive Manufacturing
- Rainbow Artifacts from Electromagnetic Signal Injection Attacks on Image Sensors
- A Hybrid Multilayer Extreme Learning Machine for Image Classification with an Application to Quadcopters
- 6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
- Bi-Classifier Determinacy Maximization for Unsupervised Domain Adaptation
- Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
- A model-agnostic active learning approach for animal detection from camera traps
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- LOVON: Legged Open-Vocabulary Object Navigator
- What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
- VisioPath: Vision-Language Enhanced Model Predictive Control for Safe Autonomous Navigation in Mixed Traffic
- SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift
- Humble Teachers Teach Better Students for Semi-Supervised Object Detection
- Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning
- Vision Based Picking System for Automatic Express Package Dispatching
- DREAM: Document Reconstruction via End-to-end Autoregressive Model
- Learning to Reconstruct and Segment 3D Objects
- SIENet: Spatial Information Enhancement Network for 3D Object Detection from Point Cloud
- DeepApple: Deep Learning-based Apple Detection using a Suppression Mask R-CNN
- Self-Knowledge Distillation with Progressive Refinement of Targets
- Batch Normalization with Enhanced Linear Transformation
- Learning Human-Object Interactions by Graph Parsing Neural Networks
- Spatio-temporal Person Retrieval via Natural Language Queries
- Actions as Moving Points
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
- YOLO-APD: Enhancing YOLOv8 for Robust Pedestrian Detection on Complex Road Geometries
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- ReFormer: The Relational Transformer for Image Captioning
- Self-Supervised Real-Time Tracking of Military Vehicles in Low-FPS UAV Footage
- Monocular 3D Object Detection and Box Fitting Trained End-to-End Using Intersection-over-Union Loss
- HGNet: High-Order Spatial Awareness Hypergraph and Multi-Scale Context Attention Network for Colorectal Polyp Detection
- Visual Identification of Individual Holstein-Friesian Cattle via Deep Metric Learning
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- LabelEnc: A New Intermediate Supervision Method for Object Detection
- Few-Shot Transformation of Common Actions into Time and Space
- DevMuT: Testing Deep Learning Framework via Developer Expertise-Based Mutation
- Automated Object Behavioral Feature Extraction for Potential Risk Analysis based on Video Sensor
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- Early Convolutions Help Transformers See Better
- ASFormer: Transformer for Action Segmentation
- Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
- Overcoming Statistical Shortcuts for Open-ended Visual Counting
- Self-supervised Object Tracking with Cycle-consistent Siamese Networks
- On Class Imbalance and Background Filtering in Visual Relationship Detection
- CUAB: Convolutional Uncertainty Attention Block Enhanced the Chest X-ray Image Analysis
- T-SYNTH: A Knowledge-Based Dataset of Synthetic Breast Images
- Unsupervised data augmentation for object detection
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- Universal Bounding Box Regression and Its Applications
- Answer Them All! Toward Universal Visual Question Answering Models
- MRC-DETR: An Adaptive Multi-Residual Coupled Transformer for Bare Board PCB Defect Detection
- R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous Driving
- C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text Spotting
- SOLQ: Segmenting Objects by Learning Queries
- Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
- Unsupervised Deep Representation Learning for Real-Time Tracking
- Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic Data
- Learning Spatio-Appearance Memory Network for High-Performance Visual Tracking
- Local Metrics for Multi-Object Tracking
- Malignancy Prediction and Lesion Identification from Clinical Dermatological Images
- CircleNet: Anchor-free Detection with Circle Representation
- Automatic Labelling for Low-Light Pedestrian Detection
- FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
- MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
- CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
- A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
- The Problem of Fragmented Occlusion in Object Detection
- DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
- Tomato plant disease detection using transfer learning with C-GAN synthetic images
- Robust Real-Time Pedestrian Detection on Embedded Devices
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- Dynamic Refinement Network for Oriented and Densely Packed Object Detection
- Towards Compact CNNs via Collaborative Compression
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
- Domain Adaptation without Source Data
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors
- Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
- Affordance Transfer Learning for Human-Object Interaction Detection
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
- SSA-CNN: Semantic Self-Attention CNN for Pedestrian Detection
- Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
- NOCTIS: Novel Object Cyclic Threshold based Instance Segmentation
- Active Measurement: Efficient Estimation at Scale
- Robust Component Detection for Flexible Manufacturing: A Deep Learning Approach to Tray-Free Object Recognition under Variable Lighting
- High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
- UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
- De-Simplifying Pseudo Labels to Enhancing Domain Adaptive Object Detection
- Disentangling Label Distribution for Long-tailed Visual Recognition
- m-RevNet: Deep Reversible Neural Networks with Momentum
- Customizable ROI-Based Deep Image Compression
- Mind the Detail: Uncovering Clinically Relevant Image Details in Accelerated MRI with Semantically Diverse Reconstructions
- Amodal Instance Segmentation
- AFAN: Augmented Feature Alignment Network for Cross-Domain Object Detection
- X-LineNet: Detecting Aircraft in Remote Sensing Images by a pair of Intersecting Line Segments
- A Revised Generative Evaluation of Visual Dialogue
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition
- SelvaBox: A high-resolution dataset for tropical tree crown detection
- Restoring Spatially-Heterogeneous Distortions using Mixture of Experts Network
- Data Augmentation For Small Object using Fast AutoAugment
- Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
- PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
- Event-based Tiny Object Detection: A Benchmark Dataset and Baseline
- Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation
- Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection
- Semi-Supervised Surface Anomaly Detection of Composite Wind Turbine Blades From Drone Imagery
- TOOD: Task-aligned One-stage Object Detection
- Data Augmentation Revisited: Rethinking the Distribution Gap between Clean and Augmented Data
- Achieving Human Parity on Visual Question Answering
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- Linear Context Transform Block
- MiniVLM: A Smaller and Faster Vision-Language Model
- OGNet: Salient Object Detection with Output-guided Attention Module
- Congestion Analysis of Convolutional Neural Network-Based Pedestrian Counting Methods on Helicopter Footage
- Person-in-WiFi: Fine-grained Person Perception using WiFi
- CAMERAS: Enhanced Resolution And Sanity preserving Class Activation Mapping for image saliency
- DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
- Dont Even Look Once: Synthesizing Features for Zero-Shot Detection
- Attention to the Burstiness in Visual Prompt Tuning!
- Region-Aware CAM: High-Resolution Weakly-Supervised Defect Segmentation via Salient Region Perception
- A Survey of Human-in-the-loop for Machine Learning
- VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
- Dual-Level Collaborative Transformer for Image Captioning
- Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation
- A New Window Loss Function for Bone Fracture Detection and Localization in X-ray Images with Point-based Annotation
- Improving Token-based Object Detection with Video
- Attention-disentangled Uniform Orthogonal Feature Space Optimization for Few-shot Object Detection
- Contextual Heterogeneous Graph Network for Human-Object Interaction Detection
- Learning Human-Object Interaction Detection using Interaction Points
- Social Adaptive Module for Weakly-supervised Group Activity Recognition
- Detecting Text in the Wild with Deep Character Embedding Network
- Disaster mapping from satellites: damage detection with crowdsourced point labels
- Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
- Object Tracking Using Spatio-Temporal Future Prediction
- Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Energy Drain of the Object Detection Processing Pipeline for Mobile Devices: Analysis and Implications
- TITAN: Query-Token based Domain Adaptive Adversarial Learning
- VideoMix: Rethinking Data Augmentation for Video Classification
- YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
- Finding Task-Relevant Features for Few-Shot Learning by Category Traversal
- Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology
- 3D for Free: Crossmodal Transfer Learning using HD Maps
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter
- LogoDet-3K: A Large-Scale Image Dataset for Logo Detection
- Class-Agnostic Region-of-Interest Matching in Document Images
- Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- LiDAR point-cloud processing based on projection methods: a comparison
- Implicit Saliency in Deep Neural Networks
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
- Dense Video Captioning using Graph-based Sentence Summarization
- Lightweight Multi-Frame Integration for Robust YOLO Object Detection in Videos
- AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing
- Feature Hallucination for Self-supervised Action Recognition
- From Codicology to Code: A Comparative Study of Transformer and YOLO-based Detectors for Layout Analysis in Historical Documents
- Multi-Scale Positive Sample Refinement for Few-Shot Object Detection
- Cascaded Regression Tracking: Towards Online Hard Distractor Discrimination
- Scalable Dynamic Origin-Destination Demand Estimation Enhanced by High-Resolution Satellite Imagery Data
- Investigating Attention Mechanism in 3D Point Cloud Object Detection
- CBAM-STN-TPS-YOLO: Enhancing Agricultural Object Detection through Spatially Adaptive Attention Mechanisms
- Geometry Normalization Networks for Accurate Scene Text Detection
- AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
- Memory Enhanced Global-Local Aggregation for Video Object Detection
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- USIS16K: High-Quality Dataset for Underwater Salient Instance Segmentation
- Cooling-Shrinking Attack: Blinding the Tracker with Imperceptible Noises
- Progressive Modality Cooperation for Multi-Modality Domain Adaptation
- Deep Reasoning with Knowledge Graph for Social Relationship Understanding
- Disentangled Motif-aware Graph Learning for Phrase Grounding
- Joint Deep Learning of Facial Expression Synthesis and Recognition
- InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
- USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
- Mask-GD Segmentation Based Robotic Grasp Detection
- Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
- Domain Randomization for Object Detection in Manufacturing Applications using Synthetic Data: A Comprehensive Study
- Spatial-temporal Fusion Convolutional Neural Network for Simulated Driving Behavior Recognition
- Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Pattern-Based Phase-Separation of Tracer and Dispersed Phase Particles in Two-Phase Defocusing Particle Tracking Velocimetry
- On the Robustness of Human-Object Interaction Detection against Distribution Shift
- Feedback Driven Multi Stereo Vision System for Real-Time Event Analysis
- Segmentation Mask Guided End-to-End Person Search
- A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & Segmentation
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- CSDN: A Context-Gated Self-Adaptive Detection Network for Real-Time Object Detection
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- HARRISON: A Benchmark on HAshtag Recommendation for Real-world Images in Social Networks
- GUIPilot: A Consistency-based Mobile GUI Testing Approach for Detecting Application-specific Bugs
- RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection
- On Feature Normalization and Data Augmentation
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- Unsupervised Pre-training for Person Re-identification
- CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
- Class Agnostic Instance-level Descriptor for Visual Instance Search
- Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
- Convolutional Character Networks
- RealDriveSim: A Realistic Multi-Modal Multi-Task Synthetic Dataset for Autonomous Driving
- End-to-End Video Instance Segmentation with Transformers
- iShape: A First Step Towards Irregular Shape Instance Segmentation
- You Cannot Easily Catch Me: A Low-Detectable Adversarial Patch for Object Detectors
- Hierarchical Scene Parsing by Weakly Supervised Learning with Image Descriptions
- Turning old models fashion again: Recycling classical CNN networks using the Lattice Transformation
- Real-Time Seamless Single Shot 6D Object Pose Prediction
- Opening up Open-World Tracking
- Translate-to-Recognize Networks for RGB-D Scene Recognition
- PointINS: Point-based Instance Segmentation
- 3D Object Detection on Point Clouds using Local Ground-aware and Adaptive Representation of scenes' surface
- AI-driven visual monitoring of industrial assembly tasks
- Image Captioning Based on a Hierarchical Attention Mechanism and Policy Gradient Optimization
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- SAFCAR: Structured Attention Fusion for Compositional Action Recognition
- Open-World Object Counting in Videos
- Vehicle Detection in Deep Learning
- Improving Visual Question Answering by Referring to Generated Paragraph Captions
- Towards Automatic Construction of Diverse, High-quality Image Dataset
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- Rethinking Pseudo-LiDAR Representation
- Learning to Navigate for Fine-grained Classification
- Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems
- GoG: Relation-aware Graph-over-Graph Network for Visual Dialog
- Malaria Detection and Classificaiton
- Multiple Instance Segmentation in Brachial Plexus Ultrasound Image Using BPMSegNet
- Egocentric Human-Object Interaction Detection: A New Benchmark and Method
- YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Hallucination Improves Few-Shot Object Detection
- Robustness Enhancement of Object Detection in Advanced Driver Assistance Systems (ADAS)
- Evaluating Explainability: A Framework for Systematic Assessment and Reporting of Explainable AI Features
- An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-Level Structural Information
- ESRPCB: an Edge guided Super-Resolution model and Ensemble learning for tiny Printed Circuit Board Defect detection
- Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
- Curriculum By Smoothing
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- ADAM-Dehaze: Adaptive Density-Aware Multi-Stage Dehazing for Improved Object Detection in Foggy Conditions
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- MeliusNet: Can Binary Neural Networks Achieve MobileNet-level Accuracy?
- Benchmarking Single Image Dehazing and Beyond
- FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception
- Bring Your Own Codegen to Deep Learning Compiler
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- Human Object Interaction Detection using Two-Direction Spatial Enhancement and Exclusive Object Prior
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Universal Lesion Detection by Learning from Multiple Heterogeneously Labeled Datasets
- Stereo Vision Based Single-Shot 6D Object Pose Estimation for Bin-Picking by a Robot Manipulator
- Attention Based Glaucoma Detection: A Large-scale Database and CNN Model
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- Improving 3D Object Detection for Pedestrians with Virtual Multi-View Synthesis Orientation Estimation
- Fingerspelling recognition in the wild with iterative visual attention
- Towards Unified INT8 Training for Convolutional Neural Network
- Generalization Bounds for Convolutional Neural Networks
- DSFD: Dual Shot Face Detector
- DMV: Visual Object Tracking via Part-level Dense Memory and Voting-based Retrieval
- Context Encoding Chest X-rays
- Affinity Graph Supervision for Visual Recognition
- Learned Image Compression for Machine Perception
- Team JL Solution to Google Landmark Recognition 2019
- Image-based localization using LSTMs for structured feature correlation
- Scene-aware SAR ship detection guided by unsupervised sea-land segmentation
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
- Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
- MatchPlant: An Open-Source Pipeline for UAV-Based Single-Plant Detection and Data Extraction
- Dynamic Relevance Learning for Few-Shot Object Detection
- Towards End-to-end Text Spotting with Convolutional Recurrent Neural Networks
- Distilling Object Detectors with Task Adaptive Regularization
- Vision-based Lifting of 2D Object Detections for Automated Driving
- Teleoperated Driving: a New Challenge for 3D Object Detection in Compressed Point Clouds
- Efficient Object Embedding for Spliced Image Retrieval
- Sequence Level Semantics Aggregation for Video Object Detection
- High-order Tensor Pooling with Attention for Action Recognition
- Graph-Structured Referring Expression Reasoning in The Wild
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Joint Layout Analysis, Character Detection and Recognition for Historical Document Digitization
- LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment
- Appearance-Preserving 3D Convolution for Video-based Person Re-identification
- Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
- UG2+ Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments
- The Multi-Modal Video Reasoning and Analyzing Competition
- A Neural Rejection System Against Universal Adversarial Perturbations in Radio Signal Classification
- EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
- ViLLa: A Neuro-Symbolic approach for Animal Monitoring
- Semantic-decoupled Spatial Partition Guided Point-supervised Oriented Object Detection
- GRIP: Generative Robust Inference and Perception for Semantic Robot Manipulation in Adversarial Environments
- Teaching in adverse scenes: a statistically feedback-driven threshold and mask adjustment teacher-student framework for object detection in UAV images under adverse scenes
- Semi-Autoregressive Image Captioning
- Deep Volumetric Universal Lesion Detection using Light-Weight Pseudo 3D Convolution and Surface Point Regression
- J-DDL: Surface Damage Detection and Localization System for Fighter Aircraft
- SLICK: Selective Localization and Instance Calibration for Knowledge-Enhanced Car Damage Segmentation in Automotive Insurance
- K-Net: Towards Unified Image Segmentation
- ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation
- Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection
- Synthesizing Unrestricted False Positive Adversarial Objects Using Generative Models
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- The Four Color Theorem for Cell Instance Segmentation
- GLD-Road:A global-local decoding road network extraction model for remote sensing images
- Exploiting the Inherent Limitation of L0 Adversarial Examples
- Learning to Collocate Neural Modules for Image Captioning
- Simple online and real-time tracking with occlusion handling
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- Dual Normalization Multitasking for Audio-Visual Sounding Object Localization
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- Dynamic Filtering with Large Sampling Field for ConvNets
- Semi-Online Knowledge Distillation
- Domain Adaptation for Big Data in Agricultural Image Analysis: A Comprehensive Review
- Adaptive Contextual Embedding for Robust Far-View Borehole Detection
- Regularizing Attention Networks for Anomaly Detection in Visual Question Answering
- VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning
- AdaPruner: Adaptive Channel Pruning and Effective Weights Inheritance
- Locality-Aware Rotated Ship Detection in High-Resolution Remote Sensing Imagery Based on Multi-Scale Convolutional Network
- Learning from Lexical Perturbations for Consistent Visual Question Answering
- Mask R-CNN with Pyramid Attention Network for Scene Text Detection
- Real-Time Privacy Preservation for Robot Visual Perception
- Context-Aware Unsupervised Clustering for Person Search
- Advancement and Field Evaluation of a Dual-arm Apple Harvesting Robot
- Online Multi-Object Tracking with Dual Matching Attention Networks
- Integer Binary-Range Alignment Neuron for Spiking Neural Networks
- Joint Visual Grounding with Language Scene Graphs
- Extract and Merge: Merging extracted humans from different images utilizing Mask R-CNN
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- HOI Analysis: Integrating and Decomposing Human-Object Interaction
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Mitosis Detection for Breast Cancer Pathology Images Using UV-Net
- FRED: The Florence RGB-Event Drone Dataset
- Bridging Annotation Gaps: Transferring Labels to Align Object Detection Datasets
- On Applying Machine Learning/Object Detection Models for Analysing Digitally Captured Physical Prototypes from Engineering Design Projects
- A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions
- VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
- Cheaper Pre-training Lunch: An Efficient Paradigm for Object Detection
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Multiple Object Tracking with Motion and Appearance Cues
- An Automatic System for Unconstrained Video-Based Face Recognition
- Learning to decompose for object detection and instance segmentation
- Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
- Image-based table recognition: data, model, and evaluation
- Multi-Branch Fully Convolutional Network for Face Detection
- Accurate Monocular Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving
- Act Like a Radiologist: Towards Reliable Multi-view Correspondence Reasoning for Mammogram Mass Detection
- Bipartite Graph Network with Adaptive Message Passing for Unbiased Scene Graph Generation
- iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images
- Object Detection in Video with Spatial-temporal Context Aggregation
- Key Points Estimation and Point Instance Segmentation Approach for Lane Detection
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- SS3D: Single Shot 3D Object Detector
- Style Normalization and Restitution for Domain Generalization and Adaptation
- Domain Contrast for Domain Adaptive Object Detection
- Exploring the Hierarchy in Relation Labels for Scene Graph Generation
- The Next Big Thing(s) in Unsupervised Machine Learning: Five Lessons from Infant Learning
- Better Understanding Hierarchical Visual Relationship for Image Caption
- Copyspace: Where to Write on Images?
- Fine-Grained Image Analysis with Deep Learning: A Survey
- Reduced Focal Loss: 1st Place Solution to xView object detection in Satellite Imagery
- Detecting Human-Object Interactions with Action Co-occurrence Priors
- Transformation Driven Visual Reasoning
- Learning Efficient Detector with Semi-supervised Adaptive Distillation
- MS-YOLO: A Multi-Scale Model for Accurate and Efficient Blood Cell Detection
- Diversifying Sample Generation for Accurate Data-Free Quantization
- Compositional Scalable Object SLAM
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait
- MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos
- 360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
- Recent Advances in Deep Learning for Object Detection
- IC Networks: Remodeling the Basic Unit for Convolutional Neural Networks
- Phrase Localization Without Paired Training Examples
- DOSED: a deep learning approach to detect multiple sleep micro-events in EEG signal
- Long-Term Feature Banks for Detailed Video Understanding
- GPS-Net: Graph Property Sensing Network for Scene Graph Generation
- Shape2Motion: Joint Analysis of Motion Parts and Attributes from 3D Shapes
- MixTConv: Mixed Temporal Convolutional Kernels for Efficient Action Recogntion
- Multi-layer Feature Aggregation for Deep Scene Parsing Models
- Weakly Supervised Cell Instance Segmentation by Propagating from Detection Response
- Meta-DETR: Image-Level Few-Shot Object Detection with Inter-Class Correlation Exploitation
- Double Anchor R-CNN for Human Detection in a Crowd
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- MaskFace: multi-task face and landmark detector
- EmbedMask: Embedding Coupling for One-stage Instance Segmentation
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- Category Anchor-Guided Unsupervised Domain Adaptation for Semantic Segmentation
- Fusion of Head and Full-Body Detectors for Multi-Object Tracking
- One-Shot Unsupervised Cross-Domain Detection
- RAPiD: Rotation-Aware People Detection in Overhead Fisheye Images
- Relation Transformer Network
- Spatial Priming for Detecting Human-Object Interactions
- Audiovisual SlowFast Networks for Video Recognition
- Down to the Last Detail: Virtual Try-on with Detail Carving
- Sequential Context Encoding for Duplicate Removal
- Cross Chest Graph for Disease Diagnosis with Structural Relational Reasoning
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- Augmented Parallel-Pyramid Net for Attention Guided Pose-Estimation
- DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
- ASK: Adaptively Selecting Key Local Features for RGB-D Scene Recognition
- Towards Accurate One-Stage Object Detection with AP-Loss
- Farm land weed detection with region-based deep convolutional neural networks
- Meta Module Network for Compositional Visual Reasoning
- Scale-Aware Trident Networks for Object Detection
- Multi-Attribute Enhancement Network for Person Search
- Adversarial Examples in Modern Machine Learning: A Review
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Goal-oriented Object Importance Estimation in On-road Driving Videos
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization
- Towards Robust Pattern Recognition: A Review
- Transformer-Based Source-Free Domain Adaptation
- MobilePose: Real-Time Pose Estimation for Unseen Objects with Weak Shape Supervision
- Towards Better Object Detection in Scale Variation with Adaptive Feature Selection
- Signet Ring Cell Detection With a Semi-supervised Learning Framework
- Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning
- DeeperLab: Single-Shot Image Parser
- Learning Joint Representations of Videos and Sentences with Web Image Search
- DiagNet: Detecting Objects using Diagonal Constraints on Adjacency Matrix of Graph Neural Network
- Adaptive Object Detection with Dual Multi-Label Prediction
- A Deep Neuro-Fuzzy Network for Image Classification
- You Only Train Once
- Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer
- Variational Saccading: Efficient Inference for Large Resolution Images
- FAN: Focused Attention Networks
- Video-Text Pre-training with Learned Regions
- A Free Lunch for Unsupervised Domain Adaptive Object Detection without Source Data
- Deeply Activated Salient Region for Instance Search
- FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
- PDSE: A Multiple Lesion Detector for CT Images using PANet and Deformable Squeeze-and-Excitation Block
- Ice Hockey Puck Localization Using Contextual Cues
- InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting
- Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- PC2WF: 3D Wireframe Reconstruction from Raw Point Clouds
- High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
- Normalized and Geometry-Aware Self-Attention Network for Image Captioning
- Reflective Decoding Network for Image Captioning
- From Multi-View to Hollow-3D: Hallucinated Hollow-3D R-CNN for 3D Object Detection
- A Parameterized Approach to Personalized Variable Length Summarization of Soccer Matches
- Rethinking Image-Scaling Attacks: The Interplay Between Vulnerabilities in Machine Learning Systems
- Weak Novel Categories without Tears: A Survey on Weak-Shot Learning
- SportMamba: Adaptive Non-Linear Multi-Object Tracking with State Space Models for Team Sports
- Geometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos
- Auto-Labeling Data for Object Detection
- Organ At Risk Segmentation with Multiple Modality
- Stacked Deconvolutional Network for Semantic Segmentation
- The More You Know: Using Knowledge Graphs for Image Classification
- Real-time Human-Robot Collaborative Manipulations of Cylindrical and Cubic Objects via Geometric Primitives and Depth Information
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
- Hierarchy Parsing for Image Captioning
- Towards Unconstrained End-to-End Text Spotting
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning
- FDSG: Forecasting Dynamic Scene Graphs
- VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
- A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
- PixelLink: Detecting Scene Text via Instance Segmentation
- WeightAlign: Normalizing Activations by Weight Alignment
- VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment
- ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos
- DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral
- OD3: Optimization-free Dataset Distillation for Object Detection
- SpatialFlow: Bridging All Tasks for Panoptic Segmentation
- Localizing Visual Sounds the Hard Way
- Optical Music Recognition: State of the Art and Major Challenges
- NODIS: Neural Ordinary Differential Scene Understanding
- SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection
- Advancing from Automated to Autonomous Beamline by Leveraging Computer Vision
- Toward Transformer-Based Object Detection
- Visual Agreement Regularized Training for Multi-Modal Machine Translation
- Visual Diagnosis of Dermatological Disorders: Human and Machine Performance
- Investigations of the Influences of a CNN's Receptive Field on Segmentation of Subnuclei of Bilateral Amygdalae
- Towards Predicting Any Human Trajectory In Context
- Pseudo-Labeling Driven Refinement of Benchmark Object Detection Datasets via Analysis of Learning Patterns
- Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
- Probabilistic Oriented Object Detection in Automotive Radar
- Feature-Driven Super-Resolution for Object Detection
- Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
- iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection
- Learning to Assemble Neural Module Tree Networks for Visual Grounding
- Object Detection Networks on Convolutional Feature Maps
- Learning Instance-wise Sparsity for Accelerating Deep Models
- Target-Tailored Source-Transformation for Scene Graph Generation
- Multi-Task Incremental Learning for Object Detection
- GTNet:Guided Transformer Network for Detecting Human-Object Interactions
- Fast Local Attack: Generating Local Adversarial Examples for Object Detectors
- Jigsaw Clustering for Unsupervised Visual Representation Learning
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Point Cloud Instance Segmentation using Probabilistic Embeddings
- Deformable Attention Mechanisms Applied to Object Detection, case of Remote Sensing
- Discriminative Semantic Feature Pyramid Network with Guided Anchoring for Logo Detection
- MFPN: A Novel Mixture Feature Pyramid Network of Multiple Architectures for Object Detection
- Error-Corrected Margin-Based Deep Cross-Modal Hashing for Facial Image Retrieval
- Discriminative Triad Matching and Reconstruction for Weakly Referring Expression Grounding
- Spectro-Temporal RF Identification using Deep Learning
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Learning to Filter: Siamese Relation Network for Robust Tracking
- Characters Detection on Namecard with faster RCNN
- Efficient Pig Counting in Crowds with Keypoints Tracking and Spatial-aware Temporal Response Filtering
- TPDNet: Triple phenotype deepen networks for monocular 3D object detection of melons and fruits in fields
- Conformal Object Detection by Sequential Risk Control
- Semi Supervised Deep Quick Instance Detection and Segmentation
- Shape Back-Projection In 3D Scenes
- PaStaNet: Toward Human Activity Knowledge Engine
- ReenactGAN: Learning to Reenact Faces via Boundary Transfer
- Image Captioning with Integrated Bottom-Up and Multi-level Residual Top-Down Attention for Game Scene Understanding
- A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
- Towards Fully Automated Manga Translation
- Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
- Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
- Advancing Image Super-resolution Techniques in Remote Sensing: A Comprehensive Survey
- WTEFNet: Real-Time Low-Light Object Detection for Advanced Driver Assistance Systems
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- Question Type Guided Attention in Visual Question Answering
- Multi-Sourced Compositional Generalization in Visual Question Answering
- SafeNet: An Assistive Solution to Assess Incoming Threats for Premises
- End-to-end Compression Towards Machine Vision: Network Architecture Design and Optimization
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- Discovering Multi-Label Actor-Action Association in a Weakly Supervised\n Setting
- Simple Unsupervised Multi-Object Tracking
- FMG-Det: Foundation Model Guided Robust Object Detection
- Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
- TinaFace: Strong but Simple Baseline for Face Detection
- Bosch Deep Learning Hardware Benchmark
- E2BoWs: An End-to-End Bag-of-Words Model via Deep Convolutional Neural Network
- Joint Monocular 3D Vehicle Detection and Tracking
- DeepIM: Deep Iterative Matching for 6D Pose Estimation
- PIDray: A Large-Scale X-ray Benchmark for Real-World Prohibited Item Detection
- A Survey on Long-Tailed Visual Recognition
- RefineMask: Towards High-Quality Instance Segmentation with Fine-Grained Features
- Face Detection with End-to-End Integration of a ConvNet and a 3D Model
- Structure Information is the Key: Self-Attention RoI Feature Extractor in 3D Object Detection
- RGBX-DiffusionDet: A Framework for Multi-Modal RGB-X Object Detection Using DiffusionDet
- Multi-Modal Pedestrian Detection with Large Misalignment Based on Modal-Wise Regression and Multi-Modal IoU
- Low-Light Maritime Image Enhancement with Regularized Illumination Optimization and Deep Noise Suppression
- Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
- Noisy Annotation Refinement for Object Detection
- Patch2Pix: Epipolar-Guided Pixel-Level Correspondences
- Soft Anchor-Point Object Detection
- Streaming Object Detection for 3-D Point Clouds
- Multi-View Multi-Person 3D Pose Estimation with Plane Sweep Stereo
- Domain-Specific Suppression for Adaptive Object Detection
- MSFNet-CPD: Multi-Scale Cross-Modal Fusion Network for Crop Pest Detection
- Large Product Key Memory for Pretrained Language Models
- Learning Natural Language Generation from Scratch
- Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers
- DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
- Semantic Feature Matching for Robust Mapping in Agriculture
- LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
- Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs
- Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
- Abstractive Text Classification Using Sequence-to-convolution Neural Networks
- SAM-RCNN: Scale-Aware Multi-Resolution Multi-Channel Pedestrian Detection
- AmphibianDetector: adaptive computation for moving objects detection
- Identifying and Exploiting Structures for Reliable Deep Learning
- Detect-to-Retrieve: Efficient Regional Aggregation for Image Search
- Open-Det: An Efficient Learning Framework for Open-Ended Detection
- Assured Autonomy with Neuro-Symbolic Perception
- YOLO-SPCI: Enhancing Remote Sensing Object Detection via Selective-Perspective-Class Integration
- Visual Product Graph: Bridging Visual Products And Composite Images For End-to-End Style Recommendations
- VisAlgae 2023: A Dataset and Challenge for Algae Detection in Microscopy Images
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Joint 3D Proposal Generation and Object Detection from View Aggregation
- Using Artificial Intelligence to Analyze Fashion Trends
- VLUC: An Empirical Benchmark for Video-Like Urban Computing on Citywide Crowd and Traffic Prediction
- Position, Padding and Predictions: A Deeper Look at Position Information in CNNs
- OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields
- UPSNet: A Unified Panoptic Segmentation Network
- Dual Attention Networks for Visual Reference Resolution in Visual Dialog
- From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
- A Novel Shape-Aware Topological Representation for GPR Data with DNN Integration
- BCNet: Searching for Network Width with Bilaterally Coupled Network
- A Contrastive Learning Foundation Model Based on Perfectly Aligned Sample Pairs for Remote Sensing Images
- SaSi: A Self-augmented and Self-interpreted Deep Learning Approach for Few-shot Cryo-ET Particle Detection
- The Role of Video Generation in Enhancing Data-Limited Action Understanding
- Adaptive and Azimuth-Aware Fusion Network of Multimodal Local Features for 3D Object Detection
- Distilling Object Detectors via Decoupled Features
- Implicit Feature Alignment: Learn to Convert Text Recognizer to Text Spotter
- Min-Entropy Latent Model for Weakly Supervised Object Detection
- Optimizing the Trade-off between Single-Stage and Two-Stage Object Detectors using Image Difficulty Prediction
- Learning a Domain Classifier Bank for Unsupervised Adaptive Object Detection
- Tasks Integrated Networks: Joint Detection and Retrieval for Image Search
- CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
- Shape-Texture Debiased Neural Network Training
- Locality-Aware Zero-Shot Human-Object Interaction Detection
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
- Revolutionizing Wildfire Detection with Convolutional Neural Networks: A VGG16 Model Approach
- CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection
- AM-LFS: AutoML for Loss Function Search
- Pillar-based Object Detection for Autonomous Driving
- Scene Text Detection with Scribble Lines
- Staircase Recognition and Location Based on Polarization Vision
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- WIDER Face and Pedestrian Challenge 2018: Methods and Results
- EMPNet: Neural Localisation and Mapping Using Embedded Memory Points
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- A Comprehensive Approach for UAV Small Object Detection with Simulation-based Transfer Learning and Adaptive Fusion
- Greedy Network Enlarging
- Federated Few-Shot Learning with Adversarial Learning
- Deep Learning Based Instance Segmentation in 3D Biomedical Images Using Weak Annotation
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- Bottom-Up Temporal Action Localization with Mutual Regularization
- Value of Temporal Dynamics Information in Driving Scene Segmentation
- Cell Detection in Microscopy Images with Deep Convolutional Neural Network and Compressed Sensing
- Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling
- EllipseNet: Anchor-Free Ellipse Detection for Automatic Cardiac Biometrics in Fetal Echocardiography
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- Image Resizing by Reconstruction from Deep Features
- LookUP: Vision-Only Real-Time Precise Underground Localisation for Autonomous Mining Vehicles
- Category-wise Attack: Transferable Adversarial Examples for Anchor Free Object Detection
- Gradient Centralization: A New Optimization Technique for Deep Neural Networks
- LANCE: Efficient Low-Precision Quantized Winograd Convolution for Neural Networks Based on Graphics Processing Units
- Towards more transferable adversarial attack in black-box manner
- Conceptual Text Region Network: Cognition-Inspired Accurate Scene Text Detection
- Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
- Image Captioning with Unseen Objects
- Convolutional Auto-encoding of Sentence Topics for Image Paragraph Generation
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
- Saccader: Improving Accuracy of Hard Attention Models for Vision
- CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
- Visual Question Answering based on Local-Scene-Aware Referring Expression Generation
- Training Compact CNNs for Image Classification using Dynamic-coded Filter Fusion
- Smart-Inspect: Micro Scale Localization and Classification of Smartphone Glass Defects for Industrial Automation
- The Devil is in Classification: A Simple Framework for Long-tail Object Detection and Instance Segmentation
- InfoDet: A Dataset for Infographic Element Detection
- How To Train Your Deep Multi-Object Tracker
- Real-time Traffic Accident Anticipation with Feature Reuse
- Self-Supervised Object Detection via Generative Image Synthesis
- A Software Architecture for Autonomous Vehicles: Team LRM-B Entry in the First CARLA Autonomous Driving Challenge
- An Accurate and Real-time Self-blast Glass Insulator Location Method Based On Faster R-CNN and U-net with Aerial Images
- MSDU-net: A Multi-Scale Dilated U-net for Blur Detection
- Towards Panoptic 3D Parsing for Single Image in the Wild
- Context-Gated Convolution
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Structured Multimodal Attentions for TextVQA
- A Holistic Approach for Data-Driven Object Cutout
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction
- 3D Aggregated Faster R-CNN for General Lesion Detection
- Task-Aware Variational Adversarial Active Learning
- Extending Dataset Pruning to Object Detection: A Variance-based Approach
- What leads to generalization of object proposals?
- DDR-ID: Dual Deep Reconstruction Networks Based Image Decomposition for Anomaly Detection
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding
- Generative Models as a Data Source for Multiview Representation Learning
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- Detailed Evaluation of Modern Machine Learning Approaches for Optic Plastics Sorting
- Boosting Weakly Supervised Object Detection with Progressive Knowledge Transfer
- AdvReal: Physical Adversarial Patch Generation Framework for Security Evaluation of Object Detection Systems
- SuperPure: Efficient Purification of Localized and Distributed Adversarial Patches via Super-Resolution GAN Models
- Self-Classification Enhancement and Correction for Weakly Supervised Object Detection
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- Object Instance Mining for Weakly Supervised Object Detection
- ULSD: Unified Line Segment Detection across Pinhole, Fisheye, and Spherical Cameras
- Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
- Learning Predicates as Functions to Enable Few-shot Scene Graph Prediction
- DP-Image: Differential Privacy for Image Data in Feature Space
- Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
- AuxDet: Auxiliary Metadata Matters for Omni-Domain Infrared Small Target Detection
- Embodied Visual Recognition
- Deeply Aligned Adaptation for Cross-domain Object Detection
- InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
- Multispectral Detection Transformer with Infrared-Centric Feature Fusion
- IA-MOT: Instance-Aware Multi-Object Tracking with Motion Consistency
- WiderPerson: A Diverse Dataset for Dense Pedestrian Detection in the Wild
- OSSDD - a New Open Dataset for Sentinel-1 Ship Detection
- DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation
- Oral Imaging for Malocclusion Issues Assessments: OMNI Dataset, Deep Learning Baselines and Benchmarking
- Self-supervised Pre-training with Hard Examples Improves Visual Representations
- Slender Object Detection: Diagnoses and Improvements
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
- Frustum Fusion: Pseudo-LiDAR and LiDAR Fusion for 3D Detection
- Training Deep Learning Based Denoisers without Ground Truth Data
- Learning to Fuse Asymmetric Feature Maps in Siamese Trackers
- A Normalized Gaussian Wasserstein Distance for Tiny Object Detection
- Robotic Monitoring of Colorimetric Leaf Sensors for Precision Agriculture
- Can artificial intelligence be integrated into pest monitoring schemes to help achieve sustainable agriculture? An entomological, management and computational perspective
- AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards
- Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning
- Learning Person Re-identification Models from Videos with Weak Supervision
- FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning
- Semantic Dense Reconstruction with Consistent Scene Segments
- Automated Quality Evaluation of Cervical Cytopathology Whole Slide Images Based on Content Analysis
- Decoupling Classifier for Boosting Few-shot Object Detection and Instance Segmentation
- RepPoints: Point Set Representation for Object Detection
- Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots
- Student Helping Teacher: Teacher Evolution via Self-Knowledge Distillation
- Switching Loss for Generalized Nucleus Detection in Histopathology
- Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
- Universal Domain Adaptation through Self Supervision
- MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation
- FPAENet: Pneumonia Detection Network Based on Feature Pyramid Attention Enhancement
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes
- AGI-Elo: How Far Are We From Mastering A Task?
- VisDA-2021 Competition Universal Domain Adaptation to Improve Performance on Out-of-Distribution Data
- DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object Detection
- Deformation Robust Roto-Scale-Translation Equivariant CNNs
- ProMi: An Efficient Prototype-Mixture Baseline for Few-Shot Segmentation with Bounding-Box Annotations
- Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
- Quantized Neural Networks via -1, +1 Encoding Decomposition and Acceleration
- KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
- A method for detecting text of arbitrary shapes in natural scenes that improves text spotting
- Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts
- FiGKD: Fine-Grained Knowledge Distillation via High-Frequency Detail Transfer
- FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
- Localize to Classify and Classify to Localize: Mutual Guidance in Object Detection
- Unpaired Image Captioning via Scene Graph Alignments
- Interpretable Neural Computation for Real-World Compositional Visual Question Answering
- Bi-Dimensional Feature Alignment for Cross-Domain Object Detection
- Training Generative Adversarial Networks in One Stage
- Single View Physical Distance Estimation using Human Pose
- Experimental Study on Automatically Assembling Custom Catering Packages With a 3-DOF Delta Robot Using Deep Learning Methods
- Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection
- Backbone Can Not be Trained at Once: Rolling Back to Pre-trained Network for Person Re-Identification
- Pix2seq: A Language Modeling Framework for Object Detection
- Adversarial Semantic Contour for Object Detection
- Depth-conditioned Dynamic Message Propagation for Monocular 3D Object Detection
- A Serverless Cloud-Fog Platform for DNN-Based Video Analytics with Incremental Learning
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Activity Driven Weakly Supervised Object Detection
- Prime-Aware Adaptive Distillation
- A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
- Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
- The Case for High-Accuracy Classification: Think Small, Think Many!
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- LIP: Learning Instance Propagation for Video Object Segmentation
- Mask R-CNN Based Object Detection for Intelligent Wireless Power Transfer
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs
- OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts
- Mixing Real and Synthetic Data to Enhance Neural Network Training -- A Review of Current Approaches
- INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
- Learnable Histogram: Statistical Context Features for Deep Neural Networks
- Self-supervised Robust Object Detectors from Partially Labelled Datasets
- Dim and Small Target Detection for Drone Broadcast Frames Based on Time-Frequency Analysis
- HWNet v2: An Efficient Word Image Representation for Handwritten Documents
- Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
- FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration
- Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking
- End-to-end Autonomous Driving Perception with Sequential Latent Representation Learning
- MLMA-Net: multi-level multi-attentional learning for multi-label object detection in textile defect images
- Learning Informative Attention Weights for Person Re-Identification
- Temporal HeartNet: Towards Human-Level Automatic Analysis of Fetal Cardiac Screening Video
- OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement
- CenterFace: Joint Face Detection and Alignment Using Face as Point
- Long Activity Video Understanding using Functional Object-Oriented Network
- Learning Feature-to-Feature Translator by Alternating Back-Propagation for Generative Zero-Shot Learning
- FSD: Feature Skyscraper Detector for Stem End and Blossom End of Navel Orange
- Balance-Oriented Focal Loss with Linear Scheduling for Anchor Free Object Detection
- HMPNet: A Feature Aggregation Architecture for Maritime Object Detection from a Shipborne Perspective
- Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
- Simple multi-dataset detection
- Robustness Analysis against Adversarial Patch Attacks in Fully Unmanned Stores
- Self-supervised learning for autonomous vehicles perception: A conciliation between analytical and learning methods
- Cloud-based Image Classification Service Is Not Robust To Simple Transformations: A Forgotten Battlefield
- Unbiased Scene Graph Generation via Rich and Fair Semantic Extraction
- Deep Reinforcement Learning for Active Human Pose Estimation
- Attribute Aware Pooling for Pedestrian Attribute Recognition
- Custom Object Detection via Multi-Camera Self-Supervised Learning
- Scene-Aware Error Modeling of LiDAR/Visual Odometry for Fusion-based Vehicle Localization
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- SparseMeXT Unlocking the Potential of Sparse Representations for HD Map Construction
- Language-Driven Dual Style Mixing for Single-Domain Generalized Object Detection
- GradAug: A New Regularization Method for Deep Neural Networks
- Autonomous Robotic Pruning in Orchards and Vineyards: a Review
- LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
- Spherical Kernel for Efficient Graph Convolution on 3D Point Clouds
- RiFCN: Recurrent Network in Fully Convolutional Network for Semantic Segmentation of High Resolution Remote Sensing Images
- AIBench: An Industry Standard Internet Service AI Benchmark Suite
- IQDet: Instance-wise Quality Distribution Sampling for Object Detection
- Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
- Transformer-Based Dual-Optical Attention Fusion Crowd Head Point Counting and Localization Network
- YOPOv2-Tracker: An End-to-End Agile Tracking and Navigation Framework from Perception to Action
- Memorizing Comprehensively to Learn Adaptively: Unsupervised Cross-Domain Person Re-ID with Multi-level Memory
- Improve Object Detection by Data Enhancement based on Generative Adversarial Nets
- Geometric-Topological Perception and Motion Prior for Real-Time Satellite Video Object Tracking
- VALISENS: A Validated Innovative Multi-Sensor System for Cooperative Automated Driving
- Feature Representation Transferring to Lightweight Models via Perception Coherence
- Aggressive Perception-Aware Navigation using Deep Optical Flow Dynamics and PixelMPC
- AutoDNNchip: An Automated DNN Chip Predictor and Builder for Both FPGAs and ASICs
- OICSR: Out-In-Channel Sparsity Regularization for Compact Deep Neural Networks
- Deep Learning for Automatic Quality Grading of Mangoes: Methods and Insights
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- The Application of Deep Learning for Lymph Node Segmentation: A Systematic Review
- SADet: Learning An Efficient and Accurate Pedestrian Detector
- Training Binary Neural Networks through Learning with Noisy Supervision
- Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection
- Automating Infrastructure Surveying: A Framework for Geometric Measurements and Compliance Assessment Using Point Cloud Data
- Progressive Localization Networks for Language-based Moment Localization
- Segmentation task for fashion and apparel
- Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid
- Anchors Based Method for Fingertips Position Estimation from a Monocular RGB Image using Deep Neural Network
- NAS-FCOS: Efficient Search for Object Detection Architectures
- White Light Specular Reflection Data Augmentation for Deep Learning Polyp Detection
- PaniCar: Securing the Perception of Advanced Driving Assistance Systems Against Emergency Vehicle Lighting
- Visual Affordance Prediction: Survey and Reproducibility
- Dynamic Graph: Learning Instance-aware Connectivity for Neural Networks
- A Simple Detector with Frame Dynamics is a Strong Tracker
- Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
- A Unified Hardware Architecture for Convolutions and Deconvolutions in CNN
- Spatio-temporal Tubelet Feature Aggregation and Object Linking in Videos
- Transforming faces into video stories -- VideoFace2.0
- QEBA: Query-Efficient Boundary-Based Blackbox Attack
- DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection under Partial Occlusion
- OMNIA Faster R-CNN: Detection in the wild through dataset merging and soft distillation
- Lattice Fusion Networks for Image Denoising
- Spatially Attentive Output Layer for Image Classification
- Road Damage Detection and Classification with Detectron2 and Faster R-CNN
- MA 3 : Model Agnostic Adversarial Augmentation for Few Shot learning
- OODTE: A Differential Testing Engine for the ONNX Optimizer
- Automated ARAT Scoring Using Multimodal Video Analysis, Multi-View Fusion, and Hierarchical Bayesian Models: A Clinician Study
- Focal-SAM: Focal Sharpness-Aware Minimization for Long-Tailed Classification
- End-to-end Flow Correlation Tracking with Spatial-temporal Attention
- Training Object Detectors from Few Weakly-Labeled and Many Unlabeled Images
- Localization Distillation for Dense Object Detection
- Farmland Parcel Delineation Using Spatio-temporal Convolutional Networks
- Lightweight Mask R-CNN for Long-Range Wireless Power Transfer Systems
- Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
- A Review on Object Pose Recovery: from 3D Bounding Box Detectors to Full 6D Pose Estimators
- Scan2Plan: Efficient Floorplan Generation from 3D Scans of Indoor Scenes
- Leveraging Localization for Multi-camera Association
- Fast Object Detection with Latticed Multi-Scale Feature Fusion
- Scale Matters: Temporal Scale Aggregation Network for Precise Action Localization in Untrimmed Videos
- Efficient Vision-based Vehicle Speed Estimation
- MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
- The Newspaper Navigator Dataset: Extracting And Analyzing Visual Content from 16 Million Historic Newspaper Pages in Chronicling America
- Self-Adaptive Reconfigurable Arrays (SARA): Using ML to Assist Scaling GEMM Acceleration
- Global Collinearity-aware Polygonizer for Polygonal Building Mapping in Remote Sensing
- PREMISE: Matching-based Prediction for Accurate Review Recommendation
- Spatial Semantic Embedding Network: Fast 3D Instance Segmentation with Deep Metric Learning
- Deep Learning Meets SAR
- Learning from Web Data with Self-Organizing Memory Module
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Recurrent Autoregressive Networks for Online Multi-Object Tracking
- Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature Sharing
- Scene Graph Generation for Better Image Captioning?
- Learning Cross-modal Context Graph for Visual Grounding
- Modularized Textual Grounding for Counterfactual Resilience
- Straight to Shapes: Real-time Detection of Encoded Shapes
- Dense Scene Multiple Object Tracking with Box-Plane Matching
- Knowledge Distillation By Sparse Representation Matching
- Self-Supervised Learning by Estimating Twin Class Distributions
- Visual Trajectory Prediction of Vessels for Inland Navigation
- Every Pixel Matters: Center-aware Feature Alignment for Domain Adaptive Object Detector
- On Robustness of Lane Detection Models to Physical-World Adversarial Attacks in Autonomous Driving
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
- MalariAI: A Label-Resilient Decoupled Framework for Annotation-Agnostic Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
- A Possible Reason for why Data-Driven Beats Theory-Driven Computer Vision
- Colonoscopy Polyp Detection: Domain Adaptation From Medical Report Images to Real-time Videos
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- Segmentation is All You Need
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Live Reconstruction of Large-Scale Dynamic Outdoor Worlds
- Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
- Multi Voxel-Point Neurons Convolution (MVPConv) for Fast and Accurate 3D Deep Learning
- mr2NST: Multi-Resolution and Multi-Reference Neural Style Transfer for Mammography
- Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
- Real-Time Grasping Strategies Using Event Camera
- A Computing Kernel for Network Binarization on PyTorch
- Scene-based Factored Attention for Image Captioning
- CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation
- Maria: A Visual Experience Powered Conversational Agent
- Zero-Shot Scene Graph Relation Prediction through Commonsense Knowledge Integration
- The Role of Artificial Intelligence in the SKA Era
- An Annotation Saved is an Annotation Earned: Using Fully Synthetic Training for Object Instance Detection
- Focal and efficient IOU loss for accurate bounding box regression
- Boosting Few-Shot Visual Learning With Self-Supervision
- Deep Texture-Aware Features for Camouflaged Object Detection
- An All-in-One Network for Dehazing and Beyond
- SynthText3D: Synthesizing Scene Text Images from 3D Virtual Worlds
- Trimmed Action Recognition, Dense-Captioning Events in Videos, and Spatio-temporal Action Localization with Focus on ActivityNet Challenge 2019
- Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
- Geometric Pose Affordance: 3D Human Pose with Scene Constraints
- Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction
- UAV maritime small object detection via frequency-spatial dual-domain synergistic perception and multi-granularity feature fusion
- Synthetic Data Generation and Adaption for Object Detection in Smart Vending Machines
- Automated Steel Bar Counting and Center Localization with Convolutional Neural Networks
- RFBTD: RFB Text Detector
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- MADAN: Multi-source Adversarial Domain Aggregation Network for Domain Adaptation
- Visual Tracking via Reliable Memories
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Adversarial Cross-Domain Action Recognition with Co-Attention
- CFUN: Combining Faster R-CNN and U-net Network for Efficient Whole Heart Segmentation
- Spiking Neural Network based Region Proposal Networks for Neuromorphic Vision Sensors
- Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network
- Fast Fine-Grained Image Classification via Weakly Supervised Discriminative Localization
- Jo-SRC: A Contrastive Approach for Combating Noisy Labels
- Horizontal-to-Vertical Video Conversion
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines
- Large-Scale Generative Data-Free Distillation
- Stereo Frustums: A Siamese Pipeline for 3D Object Detection
- Robust Attentive Deep Neural Network for Exposing GAN-generated Faces
- Mi YouTube es Su YouTube? Analyzing the Cultures using YouTube Thumbnails of Popular Videos
- Generating Question Relevant Captions to Aid Visual Question Answering
- BUZz: BUffer Zones for defending adversarial examples in image classification
- Straight to Shapes++: Real-time Instance Segmentation Made More Accurate
- STELA: A Real-Time Scene Text Detector with Learned Anchor
- EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach
- Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
- M6: A Chinese Multimodal Pretrainer
- Beyond Human Performance: A Vision-Language Multi-Agent Approach for Quality Control in Pharmaceutical Manufacturing
- Smoothed Gaussian Mixture Models for Video Classification and Recommendation
- 1st place solution for AVA-Kinetics Crossover in AcitivityNet Challenge 2020
- Multi-Channel CNN-based Object Detection for Enhanced Situation Awareness
- Multiple Object Tracking by Flowing and Fusing
- Efficient Model Performance Estimation via Feature Histories
- Deep Lidar CNN to Understand the Dynamics of Moving Vehicles
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- Conv-MPN: Convolutional Message Passing Neural Network for Structured Outdoor Architecture Reconstruction
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Make Both Ends Meet: A Synergistic Optimization Infrared Small Target Detection with Streamlined Computational Overhead
- Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
- Learning to Borrow Features for Improved Detection of Small Objects in Single-Shot Detectors
- XeMap: Contextual Referring in Large-Scale Remote Sensing Environments
- Cascade Detector Analysis and Application to Biomedical Microscopy
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
- Style-Adaptive Detection Transformer for Single-Source Domain Generalized Object Detection
- Purifying, Labeling, and Utilizing: A High-Quality Pipeline for Small Object Detection
- Pose Estimation for Non-Cooperative Rendezvous Using Neural Networks
- OG-HFYOLO :Orientation gradient guidance and heterogeneous feature fusion for deformation table cell instance segmentation
- T2ID-CAS: Diffusion Model and Class Aware Sampling to Mitigate Class Imbalance in Neck Ultrasound Anatomical Landmark Detection
- Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
- DG-DETR: Toward Domain Generalized Detection Transformer
- GLACIER: Rethinking Mass Spectrum Prediction as an Object Detection Problem
- Toward Generalizable Whole Brain Representations with High-Resolution Light-Sheet Data
- A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
- More Clear, More Flexible, More Precise: A Comprehensive Oriented Object Detection benchmark for UAV
- SynergyAmodal: Deocclude Anything with Text Control
- Crowd Detection Using Very-Fine-Resolution Satellite Imagery
- Improving Small Drone Detection Through Multi-Scale Processing and Data Augmentation
- Bridging Short- and Long-Term Dependencies: A CNN-Transformer Hybrid for Financial Time Series Forecasting
- Swapped Logit Distillation via Bi-level Teacher Alignment
- ODExAI: A Comprehensive Object Detection Explainable AI Evaluation
- CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes
- Transcending Dimensions using Generative AI: Real-Time 3D Model Generation in Augmented Reality
- Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
- AIBuildAI: An AI Agent for Automatically Building AI Models
- R-Sparse R-CNN: SAR Ship Detection Based on Background-Aware Sparse Learnable Proposals
- VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation
- Dream-Box: Object-wise Outlier Generation for Out-of-Distribution Detection
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- S3MOT: Monocular 3D Object Tracking with Selective State Space Model
- Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
- Achieving Human Parity on Visual Question Answering
- Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link Model
- WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects
- Multi-Stream Attention Learning for Monocular Vehicle Velocity and Inter-Vehicle Distance Estimation
- RetinaTrack: Online Single Stage Joint Detection and Tracking
- Design and Interpretation of Universal Adversarial Patches in Face Detection
- Improved Selective Refinement Network for Face Detection
- DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
- CompConv: A Compact Convolution Module for Efficient Feature Learning
- CS-R-FCN: Cross-supervised Learning for Large-Scale Object Detection
- The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
- AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
- A Decade of You Only Look Once (YOLO) for Object Detection: A Review
- FreqAdapt: Frequency-Adaptive Processing for RAW Object Detection
- An FPGA-Accelerated Design for Deep Learning Pedestrian Detection in Self-Driving Vehicles
- EHSOD: CAM-Guided End-to-end Hybrid-Supervised Object Detection with Cascade Refinement
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- Expressing Objects just like Words: Recurrent Visual Embedding for Image-Text Matching
- Learning Object Scale With Click Supervision for Object Detection
- Improving Open-World Object Localization by Discovering Background
- Cooperating RPN's Improve Few-Shot Object Detection
- Learning Temporal Pose Estimation from Sparsely-Labeled Videos
- ALET (Automated Labeling of Equipment and Tools): A Dataset, a Baseline and a Usecase for Tool Detection in the Wild
- A Review of 3D Object Detection with Vision-Language Models
- Multi-Grained Compositional Visual Clue Learning for Image Intent Recognition
- MASF-YOLO: An Improved YOLOv11 Network for Small Object Detection on Drone View
- Advances in Deep Learning for Hyperspectral Image Analysis—Addressing Challenges Arising in Practical Imaging Scenarios
- A Multimodal Hybrid Late-Cascade Fusion Network for Enhanced 3D Object Detection
- Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems
- CIVIL: Causal and Intuitive Visual Imitation Learning
- ShuffleBlock: Shuffle to Regularize Deep Convolutional Neural Networks
- Identifying Lead Water Service Lines Using Ultrasonic Stress Wave Propagation and 1D-Convolutional Neural Network
- Learning to Recognize Actions on Objects in Egocentric Video With Attention Dictionaries
- GoDP: Globally Optimized Dual Pathway deep network architecture for facial landmark localization in-the-wild
- Robustness Emerges Early in Training Dynamics, but Is Not Preserved
- ColorFD: A Finite-Difference Guided Black-Box Physical Adversarial Attack for Remote Sensing Object Detection
- Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
- FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening
- Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models
- Stampnet: Unsupervised Multi-Class Object Discovery
- SMOT: Single-Shot Multi Object Tracking
- Multi-Object Tracking with Siamese Track-RCNN
- Learning Fixation Point Strategy for Object Detection and Classification
- BorderDet: Border Feature for Dense Object Detection
- 3D-Aware Ellipse Prediction for Object-Based Camera Pose Estimation
- A Cost-Effective Person-Following System for Assistive Unmanned Vehicles with Deep Learning at the Edge
- Differentiable Neural Architecture Learning for Efficient Neural Network Design
- Enabling Incremental Knowledge Transfer for Object Detection at the Edge
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- SegICP: Integrated deep semantic segmentation and pose estimation
- Neural Architecture Search by Estimation of Network Structure Distributions
- Universal Lymph Node Detection in Multiparametric MRI with Selective Augmentation
- Active Perception for Ambiguous Objects Classification
- Online PCB Defect Detector On A New PCB Defect Dataset
- DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving
- Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
- Multi-Task Learning With Coarse Priors for Robust Part-Aware Person Re-Identification
- Detecting and Tracking Communal Bird Roosts in Weather Radar Data
- GraftNet: An Engineering Implementation of CNN for Fine-grained Multi-label Task
- Cell Detection from Imperfect Annotation by Pseudo Label Selection Using P-classification
- FRDet: Balanced and Lightweight Object Detector based on Fire-Residual Modules for Embedded Processor of Autonomous Driving
- SA-Net: Robust State-Action Recognition for Learning from Observations
- IAN: Combining Generative Adversarial Networks for Imaginative Face Generation
- Collaboration Analysis Using Deep Learning
- MTSGL: Multi-Task Structure Guided Learning for Robust and Interpretable SAR Aircraft Recognition
- Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
- Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning
- A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task Learning
- An attention-fused network for semantic segmentation of very-high-resolution remote sensing imagery
- Deep learning in photoacoustic tomography: current approaches and future directions
- Statistical Management of the False Discovery Rate in Medical Instance Segmentation Based on Conformal Risk Control
- Relationship Oriented Affordance Learning through Manipulation Graph Construction
- Combining Deep Learning With Optical Coherence Tomography Imaging to Determine Scalp Hair and Follicle Counts
- Driving Scene Perception Network: Real-Time Joint Detection, Depth Estimation and Semantic Segmentation
- Unsupervised Object Detection with LiDAR Clues
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- You Sense Only Once Beneath: Ultra-Light Real-Time Underwater Object Detection
- Measuring the Utilization of Public Open Spaces by Deep Learning: a Benchmark Study at the Detroit Riverfront
- Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction
- Multimodal Perception for Goal-oriented Navigation: A Survey
- The Open Images Dataset V4
- Deep-learning-based image segmentation integrated with optical microscopy for automatically searching for two-dimensional materials
- SAGA: Semantic-Aware Gray color Augmentation for Visible-to-Thermal Domain Adaptation across Multi-View Drone and Ground-Based Vision Systems
- Context Aware Grounded Teacher for Source Free Object Detection
- SuoiAI: Building a Dataset for Aquatic Invertebrates in Vietnam
- An Efficient Aerial Image Detection with Variable Receptive Fields
- EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
- Enhance Then Search: An Augmentation-Search Strategy with Foundation Models for Cross-Domain Few-Shot Object Detection
- Affinity LCFCN: Learning to Segment Fish with Weak Supervision
- Progressive Multi-Source Domain Adaptation for Personalized Facial Expression Recognition
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- Text Detection and Recognition in the Wild: A Review
- SoDA: Multi-Object Tracking with Soft Data Association
- FingerNet: Pushing The Limits of Fingerprint Recognition Using Convolutional Neural Network
- Slot Based Image Augmentation System for Object Detection
- STN-OCR: A single Neural Network for Text Detection and Text Recognition
- ISTD-YOLO: A Multi-Scale Lightweight High-Performance Infrared Small Target Detection Algorithm
- Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
- Semi-supervised Active Learning for Instance Segmentation via Scoring Predictions
- A Deep Learning-Based FPGA Function Block Detection Method With Bitstream to Image Transformation
- Driver Behavior Analysis Using Lane Departure Detection Under Challenging Conditions
- Compile Scene Graphs with Reinforcement Learning
- Recovering hard-to-find object instances by sampling context-based object proposals
- A Unified Mixture-View Framework for Unsupervised Representation Learning
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- Background Learnable Cascade for Zero-Shot Object Detection
- Captioning Images with Novel Objects via Online Vocabulary Expansion
- Object-Aware Instance Labeling for Weakly Supervised Object Detection
- HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
- Decoupled IoU Regression for Object Detection
- Autonomous and Cooperative Design of the Monitor Positions for a Team of UAVs to Maximize the Quantity and Quality of Detected Objects
- Condensing Two-stage Detection with Automatic Object Key Part Discovery
- Scaling Object Detection by Transferring Classification Weights
- Contralaterally Enhanced Networks for Thoracic Disease Detection
- Res2Net: A New Multi-Scale Backbone Architecture
- Denoising Large-Scale Image Captioning from Alt-text Data using Content\n Selection Models
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- Team PFDet's Methods for Open Images Challenge 2019
- BoxCars: Improving Fine-Grained Recognition of Vehicles Using 3-D Bounding Boxes in Traffic Surveillance
- Deep learning techniques for in-crop weed recognition in large-scale grain production systems: a review
- On the detection-to-track association for online multi-object tracking
- BiDet: An Efficient Binarized Object Detector
- Pacemaker: Intermediate Teacher Knowledge Distillation For On-The-Fly Convolutional Neural Network
- A Joint Intensity-Neuromorphic Event Imaging System for Resource Constrained Devices
- Deep divergence-based approach to clustering
- Configurable 3D Scene Synthesis and 2D Image Rendering with Per-pixel Ground Truth Using Stochastic Grammars
- Growing status observation for oil palm trees using Unmanned Aerial Vehicle (UAV) images
- ASFD: Automatic and Scalable Face Detector
- SRG: Snippet Relatedness-Based Temporal Action Proposal Generator
- Learning from Web Data: the Benefit of Unsupervised Object Localization
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- IvaNet: Learning to jointly detect and segment objets with the help of Local Top-Down Modules
- Learning a Layout Transfer Network for Context Aware Object Detection
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection
- Deep Learning for Scene Classification: A Survey
- SlowFast Networks for Video Recognition
- Learning from Counting: Leveraging Temporal Classification for Weakly Supervised Object Localization and Detection
- CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
- Context-Awareness and Interpretability of Rare Occurrences for Discovery and Formalization of Critical Failure Modes
- Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes
- Ellipse Detection and Localization with Applications to Knots in Sawn Lumber Images
- Deep learning‐based seabird detection in fisheries for seabird protection
- Structure-aware completion of photogrammetric meshes in urban road environment
- Collaborative Perception Datasets for Autonomous Driving: A Review
- Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks
- Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
- Quantum Computing Supported Adversarial Attack-Resilient Autonomous Vehicle Perception Module for Traffic Sign Classification
- Accurate Tracking of Arabidopsis Root Cortex Cell Nuclei in 3D Time-Lapse Microscopy Images Based on Genetic Algorithm
- Edge Approximation Text Detector
- Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
- MEC-Patch: Visible-Infrared Cross-Modal Adversarial Attack Driven by Intrinsic Material Emissivity Laws
- Pose-aware instance segmentation framework from cone beam CT images for tooth segmentation
- Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
- Graph Network for Sign Language Tasks
- PatrolVision: Automated License Plate Recognition in the wild
- CFIS-YOLO: A Lightweight Multi-Scale Fusion Network for Edge-Deployable Wood Defect Detection
- YOLO-RS: Remote Sensing Enhanced Crop Detection Methods
- GATE3D: Generalized Attention-based Task-synergized Estimation in 3D*
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- Density-based Object Detection in Crowded Scenes
- Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
- Small Object Detection with YOLO: A Performance Analysis Across Model Versions and Hardware
- Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
- A survey on recent trends in robotics and artificial intelligence in the furniture industry
- Using Vision Language Models for Safety Hazard Identification in Construction
- MBE-ARI: A Multimodal Dataset Mapping Bi-directional Engagement in Animal-Robot Interaction
- Automatic recognition of isolated piglet outliers based on multi-object tracking
- VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
- Title block detection and information extraction for enhanced building drawings search
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
- pychop: Emulating Low-Precision Arithmetic in Numerical Methods and Neural Networks
- RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
- WS-DETR: Robust Water Surface Object Detection through Vision-Radar Fusion with Detection Transformer
- A Scale Invariant Flatness Measure for Deep Network Minima
- Visually Similar Pair Alignment for Robust Cross-Domain Object Detection
- Class Imbalance Correction for Improved Universal Lesion Detection and Tagging in CT
- D-Feat Occlusions: Diffusion Features for Robustness to Partial Visual Occlusions in Object Recognition
- Joint Learning of Neural Transfer and Architecture Adaptation for Image Recognition
- Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
- GAMDTP: Dynamic Trajectory Prediction with Graph Attention Mamba Network
- Feedback-Enhanced Hallucination-Resistant Vision-Language Model for Real-Time Scene Understanding
- Strawberry Detection using Mixed Training on Simulated and Real Data
- DebGCD: Debiased Learning with Distribution Guidance for Generalized Category Discovery
Related