End-to-End Object Detection with Transformers
2020/05/26 by Nicolas Carion, Francisco Massa, Carion, Nicolas +9 · 1 voice · 508 citations
#cs.CV
paper · pdf · doi:10.48550/arxiv.2005.12872
Abstract
We present a new method that views object detection as a direct set prediction problem. Our approach streamlines the detection pipeline, effectively removing the need for many hand-designed components like a non-maximum suppression procedure or anchor generation that explicitly encode our prior knowledge about the task. The main ingredients of the new framework, called DEtection TRansformer or DETR, are a set-based global loss that forces unique predictions via bipartite matching, and a transformer encoder-decoder architecture. Given a fixed small set of learned object queries, DETR reasons about the relations of the objects and the global image context to directly output the final set of predictions in parallel. The new model is conceptually simple and does not require a specialized library, unlike many other modern detectors. DETR demonstrates accuracy and run-time performance on par with the well-established and highly-optimized Faster RCNN baseline on the challenging COCO object detection dataset. Moreover, DETR can be easily generalized to produce panoptic segmentation in a unified manner. We show that it significantly outperforms competitive baselines. Training code and pretrained models are available at https://github.com/facebookresearch/detr.
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- DeFloMat: Detection with Flow Matching for Stable and Efficient Generative Object Localization
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
- Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
- Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
- Small-Pollinator Detection in Cluttered Field Video
- Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
- CellMamba: Adaptive Mamba for Accurate and Efficient Cell Detection
- Vision Transformers are Circulant Attention Learners
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- milliMamba: Specular-Aware Human Pose Estimation via Dual mmWave Radar with Multi-Frame Mamba Fusion
- Progressive Learned Image Compression for Machine Perception
- PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
- Multi-Part Object Representations via Graph Structures and Co-Part Discovery
- Object-Centric Framework for Video Moment Retrieval
- YolovN-CBi: A Lightweight and Efficient Architecture for Real-Time Detection of Small UAVs
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- ALIGN: Advanced Query Initialization with LiDAR-Image Guidance for Occlusion-Robust 3D Object Detection
- A two-stream network with global-local feature fusion for bone age assessment
- Name That Part: 3D Part Segmentation and Naming
- StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
- Memory-Enhanced SAM3 for Occlusion-Robust Surgical Instrument Segmentation
- Tiny Recursive Control: Iterative Reasoning for Efficient Optimal Control
- DenseBEV: Transforming BEV Grid Cells into 3D Objects
- FlowDet: Unifying Object Detection and Generative Transport Flows
- OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
- PoseMoE: Mixture-of-Experts Network for Monocular 3D Human Pose Estimation
- Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
- Auto-Vocabulary 3D Object Detection
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- Particulate: Feed-Forward 3D Object Articulation
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders
- MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- Generative Spatiotemporal Data Augmentation
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- 3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
- Stronger Normalization-Free Transformers
- PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- Optimal transport unlocks end-to-end learning for single-molecule localization
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
- Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
- Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- sim2art: Accurate Articulated Object Modeling from a Single Video using Synthetic Training Data Only
- DFIR-DETR: Frequency Domain Enhancement and Dynamic Feature Aggregation for Cross-Scene Small Object Detection
- Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- TextMamba: Scene Text Detector with Mamba
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- Fast SceneScript: Accurate and Efficient Structured Language Model via Multi-Token Prediction
- VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
- KH-FUNSD: A Hierarchical and Fine-Grained Layout Analysis Dataset for Low-Resource Khmer Business Document
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- CaFTRA: Frequency-Domain Correlation-Aware Feedback-Free MIMO Transmission and Resource Allocation for 6G and Beyond
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- ViT3: Unlocking Test-Time Training in Vision
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- TBT-Former: Learning Temporal Boundary Distributions for Action Localization
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Generalized Medical Phrase Grounding
- Towards aligned body representations in vision models
- CogEvo-Edu: Cognitive Evolution Educational Multi-Agent Collaborative System
- ReactionMamba: Generating Short &Long Human Reaction Sequences
- Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
- DM3T: Harmonizing Modalities via Diffusion for Multi-Object Tracking
- Hierarchical Feature Integration for Multi-Signal Automatic Modulation Recognition
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Autonomous labeling of surgical resection margins using a foundation model
- Referring Video Object Segmentation with Cross-Modality Proxy Queries
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
- Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
- RefTr: Recurrent Refinement of Confluent Trajectories for 3D Vascular Tree Centerline Graphs
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
- AudioScene: Integrating Object-Event Audio into 3D Scenes
- SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- Exploring State-of-the-art models for Early Detection of Forest Fires
- Intelligent Image Search Algorithms Fusing Visual Large Models
- HybriDLA: Hybrid Generation for Document Layout Analysis
- Medal S: Spatio-Textual Prompt Model for Medical Segmentation
- DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
- From Features to Reference Points: Lightweight and Adaptive Fusion for Cooperative Autonomous Driving
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Exploring Surround-View Fisheye Camera 3D Object Detection
- Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cache
- Robust Physical Adversarial Patches Using Dynamically Optimized Clusters
- LRDUN: A Low-Rank Deep Unfolding Network for Efficient Spectral Compressive Imaging
- A Tri-Modal Dataset and a Baseline System for Tracking Unmanned Aerial Vehicles
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- REXO: Indoor Multi-View Radar Object Detection via 3D Bounding Box Diffusion
- RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
- Pillar-0: A New Frontier for Radiology Foundation Models
- Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
- T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving
- Benchmarking Table Extraction from Heterogeneous Scientific Extraction Documents
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
- IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
- Fast Post-Hoc Confidence Fusion for 3-Class Open-Set Aerial Object Detection
- Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
- Graph Query Networks for Object Detection with Automotive Radar
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
- MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
- Towards Metric-Aware Multi-Person Mesh Recovery by Jointly Optimizing Human Crowd in Camera Space
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- NOA: a versatile, extensible tool for AI-based organoid analysis
- MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- Fine-Grained Representation for Lane Topology Reasoning
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- SOTFormer: A Minimal Transformer for Unified Object Tracking and Trajectory Prediction
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI
- Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- Generalized-Scale Object Counting with Gradual Query Aggregation
- WEDepth: Efficient Adaptation of World Knowledge for Monocular Depth Estimation
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Visual Bridge: Universal Visual Perception Representations Generating
- Fast Multi-Organ Fine Segmentation in CT Images with Hierarchical Sparse Sampling and Residual Transformer
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
- LeCoT: revisiting network architecture for two-view correspondence pruning
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- Seq2Seq Models Reconstruct Visual Jigsaw Puzzles without Seeing Them
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
- Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
- HideAndSeg: an AI-based tool with automated prompting for octopus segmentation in natural habitats
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
- MIQ-SAM3D: From Single-Point Prompt to Multi-Instance Segmentation via Competitive Query Refinement
- MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
- Zero-Shot Multi-Animal Tracking in the Wild
- The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- RIS-Assisted 3D Spherical Splatting for Object Composition Visualization using Detection Transformers
- 3EED: Ground Everything Everywhere in 3D
- CGF-DETR: Cross-Gated Fusion DETR for Enhanced Pneumonia Detection in Chest X-rays
- Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- GDROS: A Geometry-Guided Dense Registration Framework for Optical-SAR Images under Large Geometric Transformations
- Sewer pipeline condition assessment and defect detection using computer vision
- Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
- Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features
- Detecting Unauthorized Vehicles using Deep Learning for Smart Cities: A Case Study on Bangladesh
- Image Quality Dependent Degradation for AI Systems
- DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
- Test-Time Adaptive Object Detection with Foundation Model
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
- Exponential Dynamic Energy Network for High Capacity Sequence Memory
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- MIC-BEV: Multi-Infrastructure Camera Bird's-Eye-View Transformer with Relation-Aware Fusion for 3D Object Detection
- GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- DAMap: Distance-aware MapNet for High Quality HD Map Construction
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking
- Augmenting Moment Retrieval: Zero-Dependency Two-Stage Learning
- GRASPLAT: Enabling dexterous grasping through novel view synthesis
- Kinematic Analysis and Integration of Vision Algorithms for a Mobile Manipulator Employed Inside a Self-Driving Laboratory
- A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Learning Human-Object Interaction as Groups
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers
- Towards 3D Objectness Learning in an Open World
- Closed-Loop Transfer for Weakly-supervised Affordance Grounding
- An empirical study of the effect of video encoders on Temporal Video Grounding
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- Decorrelation Speeds Up Vision Transformers
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- UniVector: Unified Vector Extraction via Instance-Geometry Interaction
- Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
- Detect Anything via Next Point Prediction
- SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding
- MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking
- CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
- Multi-Action Self-Improvement for Neural Combinatorial Optimization
- Source-Free Object Detection with Detection Transformer
- Visual Odometry with Transformers
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- YOLOv11-Litchi: Efficient Litchi Fruit Detection based on UAV-Captured Agricultural Imagery in Complex Orchard Environments
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Scaling Traffic Insights with AI and Language Model-Powered Camera Systems for Data-Driven Transportation Decision Making
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- ClustViT: Clustering-based Token Merging for Semantic Segmentation
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- Holistic Order Prediction in Natural Scenes
- FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
- Utilizing dynamic sparsity on pretrained DETR
- TARO: Toward Semantically Rich Open-World Object Detection
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
- Curriculum Learning with Synthetic Data for Enhanced Pulmonary Nodule Detection in Chest Radiographs
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- Consistent Assistant Domains Transformer for Source-free Domain Adaptation
- Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
- Vi-TacMan: Articulated Object Manipulation via Vision and Touch
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
- Cross-View Open-Vocabulary Object Detection in Aerial Imagery
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- Referring Expression Comprehension for Small Objects
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Forestpest-YOLO: A High-Performance Detection Framework for Small Forestry Pests
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
- Causally Guided Gaussian Perturbations for Out-Of-Distribution Generalization in Medical Imaging
- The Impact of Scaling Training Data on Adversarial Robustness
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
- INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative Perception
- Sim-DETR: Unlock DETR for Temporal Sentence Grounding
- A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
- From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- Motion-Aware Transformer for Multi-Object Tracking
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation
- Think with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain
- Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Edge Prediction for Roof Wireframe Reconstruction with Transformers
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- RiO-DETR: DETR for Real-time Oriented Object Detection
- CoWTracker: Tracking by Warping instead of Correlation
- DroneKey: Drone 3D Pose Estimation in Image Sequences using Gated Key-representation and Pose-adaptive Learning
- Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
- Knowledge Transfer from Interaction Learning
- SynapFlow: A Modular Framework Towards Large-Scale Analysis of Dendritic Spines
- BiGraspFormer: End-to-End Bimanual Grasp Transformer
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
- CGTGait: Collaborative Graph and Transformer for Gait Emotion Recognition
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Sparse Multiview Open-Vocabulary 3D Detection
- Region-Aware Deformable Convolutions
- ORCA: Agentic Reasoning For Hallucination and Adversarial Robustness in Vision-Language Models
- BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
- Data Leakage in Visual Datasets
- VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement
- White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Mitigating Query Selection Bias in Referring Video Object Segmentation
- Ensemble of Pre-Trained Models for Long-Tailed Trajectory Prediction
- SAGA: Selective Adaptive Gating for Efficient and Expressive Linear Attention
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew
- Advancing Real-World Parking Slot Detection with Large-Scale Dataset and Semi-Supervised Baseline
- T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
- Contextualized Representation Learning for Effective Human-Object Interaction Detection
- Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
- Multi Anatomy X-Ray Foundation Model
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- RailSafeNet: Visual Scene Understanding for Tram Safety
- Advanced Layout Analysis Models for Docling
- WebSight: A Vision-First Architecture for Robust Web Agents
- I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- Towards Understanding Visual Grounding in Visual Language Models
- mRadNet: A Compact Radar Object Detector with MetaFormer
- Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
- Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection
- RT-DETR++ for UAV Object Detection
- FPI-Det: a face--phone Interaction Dataset for phone-use detection and understanding
- IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection
- Improvement of Human-Object Interaction Action Recognition Using Scene Information and Multi-Task Learning Approach
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- ArgoTweak: Towards Self-Updating HD Maps through Structured Priors
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- Transformer-Based Neural Network for Transient Detection without Image Subtraction
- CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
- A Dataset and Benchmark for Robotic Cloth Unfolding Grasp Selection: The ICRA 2024 Cloth Competition
- UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
- Integrated Detection and Tracking Based on Radar Range-Doppler Feature
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Pointing-Guided Target Estimation via Transformer-Based Attention
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
- Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
- SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
- Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather
- Heatmap Guided Query Transformers for Robust Astrocyte Detection across Immunostains and Resolutions
- Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
- Anisotropic Fourier Features for Positional Encoding in Medical Imaging
- Vision encoders should be image size agnostic and task driven
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- NOOUGAT: Towards Unified Online and Offline Multi-Object Tracking
- An Investigation of Visual Foundation Models Robustness
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- RT-DETRv2 Explained in 8 Illustrations
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
- An End-to-End Framework for Video Multi-Person Pose Estimation
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Challenges in Non-Polymeric Crystal Structure Prediction: Why a Geometric, Permutation-Invariant Loss is Needed
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- Double-Constraint Diffusion Model with Nuclear Regularization for Ultra-low-dose PET Reconstruction
- DriveQA: Passing the Driving Knowledge Test
- Mapping like a Skeptic: Probabilistic BEV Projection for Online HD Mapping
- How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
- Representation Learning with Adaptive Superpixel Coding
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- End-to-End Analysis of Charge Stability Diagrams with Transformers
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
- Self-supervised structured object representation learning
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- POEv2: a flexible and robust framework for generic line segment detection and wireframe line segment detection
- FreeVPS: Repurposing Training-Free SAM2 for Generalizable Video Polyp Segmentation
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Weed Detection in Challenging Field Conditions: A Semi-Supervised Framework for Overcoming Shadow Bias and Data Scarcity
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- Autoregressive Universal Video Segmentation Model
- Aligning Moments in Time using Video Queries
- DQEN: Dual Query Enhancement Network for DETR-based HOI Detection
- Rethinking Human-Object Interaction Evaluation for both Vision-Language Models and HOI-Specific Methods
- Quantitative Outcome-Oriented Assessment of Microsurgical Anastomosis
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring
- You Only Pose Once: A Minimalist's Detection Transformer for Monocular RGB Category-level 9D Multi-Object Pose Estimation
- Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels
- Incremental Object Detection with Prompt-based Methods
- Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
- Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
- GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering
- Fracture Detection and Localisation in Wrist and Hand Radiographs using Detection Transformer Variants
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Wavy Transformer
- SocialTrack: Multi-Object Tracking in Complex Urban Traffic Scenes Inspired by Social Behavior
- Real-Time Beach Litter Detection and Counting: A Comparative Analysis of RT-DETR Model Variants
- GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
- TASER: Table Agents for Schema-guided Extraction and Recommendation
- Data Shift of Object Detection in Autonomous Driving
- ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages
- Controlling Multimodal LLMs via Reward-guided Decoding
- MultiPark: Multimodal Parking Transformer with Next-Segment Prediction
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- HyperTea: A Hypergraph-based Temporal Enhancement and Alignment Network for Moving Infrared Small Target Detection
- CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- MPT: Motion Prompt Tuning for Micro-Expression Recognition
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Prompt-Guided Relational Reasoning for Social Behavior Understanding with Vision Foundation Models
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- Leveraging GNN to Enhance MEF Method in Predicting ENSO
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- NS-FPN: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
- Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
- Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Physical Adversarial Camouflage through Gradient Calibration and Regularization
- Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- X-SAM: From Segment Anything to Any Segmentation
- Drone Detection with Event Cameras
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- Length Matters: Length-Aware Transformer for Temporal Sentence Grounding
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
- Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
- AttZoom: Attention Zoom for Better Visual Features
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
- MVTOP: Multi-View Transformer-based Object Pose-Estimation
- Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion
- IAUNet: Instance-Aware U-Net
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- RMT-PPAD: Real-time Multi-task Learning for Panoptic Perception in Autonomous Driving
- A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Representation Shift: Unifying Token Compression with FlashAttention
- CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective
- Multimodal Referring Segmentation: A Survey
- Stable at Any Speed: Speed-Driven Multi-Object Tracking with Learnable Kalman Filtering
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- PriorFusion: Unified Integration of Priors for Robust Road Perception in Autonomous Driving
Discussions
Related