Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
2021/10/01 by Ze Liu, Yutong Lin, Yue Cao +5 · 743 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
paper · doi:10.1109/iccv48922.2021.00986
openalex publication_date 2021/10/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at https://github.com/microsoft/Swin-Transformer.
Citations
Cited by
- A ConvNet for the 2020s
- MetaFormer is Actually What You Need for Vision
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- <scp>CasUNeXt</scp>: A Cascaded Transformer With Intra‐ and Inter‐Scale Information for Medical Image Segmentation
- DBIA: Data-free Backdoor Injection Attack against Transformer Networks
- Acoustic source localization by deep-learning attention-based modulation of microphone array data
- Automatic Quantification of Serial PET/CT Images for Pediatric Hodgkin Lymphoma Using a Longitudinally Aware Segmentation Network
- SNRAware: Improved Deep Learning MRI Denoising with Signal-to-Noise Ratio Unit Training and G-Factor Map Augmentation
- Uformer: A General U-Shaped Transformer for Image Restoration
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- VOLO: Vision Outlooker for Visual Recognition
- A multimodal whole-slide foundation model for pathology
- BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens
- A convolutional-transformer reinforcement learning agent for rotating machinery fault diagnosis
- ResT: An Efficient Transformer for Visual Recognition
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- nnFormer: Volumetric Medical Image Segmentation via a 3D Transformer
- TOPIQ: A Top-Down Approach From Semantics to Distortions for Image Quality Assessment
- Tree semantic segmentation from aerial image time series
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Fuzzy-ViT: A Deep Neuro-Fuzzy System for Cross-Domain Transfer Learning From Large-Scale General Data to Medical Image
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- Boosting Salient Object Detection with Transformer-based Asymmetric Bilateral U-Net
- The Robust Semantic Segmentation UNCV2023 Challenge Results
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- Recognition of European mammals and birds in camera trap images using deep neural networks
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- MVT: Multi-view Vision Transformer for 3D Object Recognition
- ViT-Transformer: Self-attention mechanism based constitutive modeling for nonlinear heterogeneous materials
- Attentive multilayer fusion for vision transformers
- S2-MLP: Spatial-Shift MLP Architecture for Vision
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Energy-Efficient Autonomous Driving with Adaptive Perception and Robust Decision
- Test-Time Adaptive Object Detection with Foundation Model
- Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- DRIP: Dynamic patch Reduction via Interpretable Pooling
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Decoupling What to Count and Where to See for Referring Expression Counting
- Unlocking Out-of-Distribution Generalization in Dynamics through Physics-Guided Augmentation
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- UniField: Joint Multi-Domain Training for Universal Surface Pressure Modeling
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- SARATR-X: Toward Building a Foundation Model for SAR Target Recognition
- Anomaly Detection for Medical Images Using Heterogeneous Auto-Encoder
- FocalTransNet: A Hybrid Focal-Enhanced Transformer Network for Medical Image Segmentation
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- Revealing the Potential of Learnable Perturbation Ensemble Forecast Model for Tropical Cyclone Prediction
- A Survey on Efficient Vision-Language-Action Models
- 3rd Place Scheme on Instance Segmentation Track of ICCV 2021 VIPriors Challenges
- Provable test-time adaptivity and distributional robustness of in-context learning
- Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Transforming volcanic monitoring: A dataset and benchmark for onboard volcano activity detection
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- DAMap: Distance-aware MapNet for High Quality HD Map Construction
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- PSScreen V2: Partially Supervised Multiple Retinal Disease Screening
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- LO-SDA: Latent Optimization for Score-based Atmospheric Data Assimilation
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- Efficient Large-Deformation Medical Image Registration via Recurrent Dynamic Correlation
- Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
- Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You Need
- Simplifying Knowledge Transfer in Pretrained Models
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- BCRA: bidirectional cross-modal implicit relation reasoning and aligning for text-to-image person retrieval
- Attention-guided few-shot learning for metal surface defect classification
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- FrameShield: Adversarially Robust Video Anomaly Detection
- AutoOpt: A Dataset and a Unified Framework for Automating Optimization Problem Solving
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Relieving the Over-Aggregating Effect in Graph Transformers
- LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
- Memory Constrained Dynamic Subnetwork Update for Transfer Learning
- Focal Modulation and Bidirectional Feature Fusion Network for Medical Image Segmentation
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- SutureBot: A Precision Framework & Benchmark For Autonomous End-to-End Suturing
- Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
- FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking
- Guiding diffusion models to reconstruct flow fields from sparse data
- Study of Training Dynamics for Memory-Constrained Fine-Tuning
- DARE: A Deformable Adaptive Regularization Estimator for Learning-Based Medical Image Registration
- Seabed-Net: A multi-task network for joint bathymetry estimation and seabed classification from remote sensing imagery in shallow waters
- SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
- AegisRF: Adversarial Perturbations Guided with Sensitivity for Protecting Intellectual Property of Neural Radiance Fields
- Matrix-Free Least Squares Solvers: Values, Gradients, and What to Do With Them
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- Detection and Simulation of Urban Heat Islands Using a Fine-Tuned Geospatial Foundation Model for Microclimate Impact Prediction
- A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- ProLAP: Probabilistic Language-Audio Pre-Training
- MCANet: A Coherent Multimodal Collaborative Attention Network for Advanced Modulation Recognition in Adverse Noisy Environments
- Integrated representational signatures strengthen specificity in brains and models
- Δt-Mamba3D: A Time-Aware Spatio-Temporal State-Space Model for Breast Cancer Risk Prediction
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- Rethinking PCA Through Duality
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Practical guidelines for cell segmentation models under optical aberrations in microscopy
- Facial Expression-based Parkinson's Disease Severity Diagnosis via Feature Fusion and Adaptive Class Balancing
- M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- Confidence-Weighted Semi-Supervised Learning for Skin Lesion Segmentation Using Hybrid CNN-Transformer Networks
- BARL: Bilateral Alignment in Representation and Label Spaces for Semi-Supervised Volumetric Medical Image Segmentation
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Symmetric Entropy-Constrained Video Coding for Machines
- Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
- UKANFormer: Noise-Robust Semantic Segmentation for Coral Reef Mapping via a Kolmogorov-Arnold Network-Transformer Hybrid
- CARDIUM: Congenital Anomaly Recognition with Diagnostic Images and Unified Medical records
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Video Swin Transformer
- Cost Savings from Automatic Quality Assessment of Generated Images
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- EuroMineNet: A Multitemporal Sentinel-2 Benchmark for Spatiotemporal Mining Footprint Analysis in the European Union (2015-2024)
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
- DRBD-Mamba for Robust and Efficient Brain Tumor Segmentation with Analytical Insights
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- LOTA: Bit-Planes Guided AI-Generated Image Detection
- Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- EfficientPhys: Enabling Simple, Fast and Accurate Camera-Based Vitals Measurement
- nnFormer: Interleaved Transformer for Volumetric Segmentation
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
- MAPS: Masked Attribution-based Probing of Strategies- A computational framework to align human and model explanations
- Chimera: State Space Models Beyond Sequences
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- R-Drop: Regularized Dropout for Neural Networks
- High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network
- Ripple Transformer: A Human-Object Interaction Backbone and a New Prediction Strategy for Smart Surveillance Devices
- Exploring and Leveraging Class Vectors for Classifier Editing
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- Source-Free Object Detection with Detection Transformer
- MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
- Learning Model Representations Using Publicly Available Model Hubs
- VGDM: Vision-Guided Diffusion Model for Brain Tumor Detection and Segmentation
- Deep Cross-Branch Multi-Modal Fusion Network for early Alzheimer’s diagnosis
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- LocalViT: Analyzing Locality in Vision Transformers
- A Large-Scale Benchmark for Food Image Segmentation
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks
- TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading
- SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
- Holistic Order Prediction in Natural Scenes
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- 3D Reconstruction from Transient Measurements with Time-Resolved Transformer
- Efficient Resource-Constrained Training of Vision Transformers via Subspace Optimization
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
- Vision Language Models: A Survey of 26K Papers
- VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
- SatFusion: A Unified Framework for Enhancing Satellite IoT Images via Multi-Temporal and Multi-Source Data Fusion
- Robust Canonicalization through Bootstrapped Data Re-Alignment
- SkipSR: Faster Super Resolution with Token Skipping
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- NNDM: NNUNet Diffusion Model for Brain Tumor Segmentation
- Knowledge-Aware Mamba for Joint Change Detection and Classification from MODIS Times Series
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Automated Neural Architecture Design for Industrial Defect Detection
- Spatial Uncertainty Quantification in Wildfire Forecasting for Climate-Resilient Emergency Planning
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
- HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
- Lung Infection Severity Prediction Using Transformers with Conditional TransMix Augmentation and Cross-Attention
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- TransFIRA: Transfer Learning for Face Image Recognizability Assessment
- Zeeman: A Deep Learning Regional Atmospheric Chemistry Transport Model
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Critical attention scaling in long-context transformers
- Human Action Recognition from Point Clouds over Time
- A Total Variation Regularized Framework for Epilepsy-Related MRI Image Segmentation
- Diffusion2: Turning 3D Environments into Radio Frequency Heatmaps
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- TinyViT-Batten: Few-Shot Vision Transformer with Explainable Attention for Early Batten-Disease Detection on Pediatric MRI
- Detection of retinal diseases using an accelerated reused convolutional network
- Learning more physically realistic dynamics in machine-learning based weather forecasting with latent-space constraints
- ReTiDe: Real-Time Denoising for Energy-Efficient Motion Picture Processing with FPGAs
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
- MambaCAFU: Hybrid Multi-Scale and Multi-Attention Model with Mamba-Based Fusion for Medical Image Segmentation
- Allocation of Parameters in Transformers
- Referring Expression Comprehension for Small Objects
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- CVSM: Contrastive Vocal Similarity Modeling
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Image Generation Based on Image Style Extraction
- TextCAM: Explaining Class Activation Map with Text
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
- Uformer: A General U-Shaped Transformer for Image Restoration
- SimMIM: A Simple Framework for Masked Image Modeling
- DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
- Remote Auditing: Design-based Tests of Randomization, Selection, and Missingness with Broadly Accessible Satellite Imagery
- Data driven approaches in nanophotonics: A review of AI-enabled metadevices
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- Transformer Classification of Breast Lesions: The BreastDCEDLAMBL Benchmark Dataset and 0.92 AUC Baseline
- PRISM: Progressive Rain removal with Integrated State-space Modeling
- AttriGen: Automated Multi-Attribute Annotation for Blood Cell Datasets
- Indirect Attention: Turning Context Misalignment into a Feature
- VRWKV-Editor: Reducing quadratic complexity in transformer-based video editing
- The Impact of Scaling Training Data on Adversarial Robustness
- ProbMed: A Probabilistic Framework for Medical Multimodal Binding
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Bayesian Transformer for Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
- LayerD: Decomposing Raster Graphic Designs into Layers
- Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
- BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- DRCP: Diffusion on Reinforced Cooperative Perception for Perceiving Beyond Limits
- Accurate Cobb Angle Estimation via SVD-Based Curve Detection and Vertebral Wedging Quantification
- DRIFT-Net: A Spectral--Coupled Neural Operator for PDEs Learning
- Mask Clustering-based Annotation Engine for Large-Scale Submeter Land Cover Mapping
- An Enhanced Pyramid Feature Network Based on Long-Range Dependencies for Multi-Organ Medical Image Segmentation
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- FSDENet: A Frequency and Spatial Domains based Detail Enhancement Network for Remote Sensing Semantic Segmentation
- BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
- An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
- Variable Rate Image Compression via N-Gram Context based Swin-transformer
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
- Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Enhanced Fracture Diagnosis Based on Critical Regional and Scale Aware in YOLO
- Graph Your Own Prompt
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Adversarial Scheduling
- Automated and highly precise surface wetting contact angle measurement with optical coherence tomography based on deep learning model
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer
- Seeing Isn't Believing: Context-Aware Adversarial Patch Synthesis via Conditional GAN
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- Low-cost, autonomous microscopy using deep learning and robotics: A crystal morphology case study
- Orochi: Versatile Biomedical Image Processor
- CCNeXt: An Effective Self-Supervised Stereo Depth Estimation Approach
- Category Discovery: An Open-World Perspective
- Integrating Background Knowledge in Medical Semantic Segmentation with Logic Tensor Networks
- Deep Learning-Based Cross-Anatomy CT Synthesis Using Adapted nnResU-Net with Anatomical Feature Prioritized Loss
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining
- Motion-Aware Transformer for Multi-Object Tracking
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
- Punching Above Precision: Small Quantized Model Distillation with Learnable Regularizer
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- WDformer: A Wavelet-based Differential Transformer Model for Time Series Forecasting
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
- Efficient Self-supervised Vision Transformers for Representation Learning
- HiPerformer: A High-Performance Global-Local Segmentation Model with Modular Hierarchical Fusion Strategy
- Downscaling climate projections to 1 km with single-image super resolution
- Does the Manipulation Process Matter? RITA: Reasoning Composite Image Manipulations via Reversely-Ordered Incremental-Transition Autoregression
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- Timeliness-Aware Joint Source and Channel Coding for Adaptive Image Transmission
- Towards Self-Supervised Foundation Models for Critical Care Time Series
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Myosotis: structured computation for attention like layer
- Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
- MTKGR: multi-task knowledge graph reasoning for food and ingredient recognition
- What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
- CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography
- Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
- A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- Hybrid Quantum-MambaVision: A Quantum-Enhanced State Space Model for Calibrated Mixed-Type Wafer Defect Detection
- Local and Global Feature-Aware Dual-Branch Networks for Plant Disease Recognition
- Engineered E. coli swarming for binary and analog input recording
- LAAP: Learning the Argument of An Entity with Event Prompts for document-level event extraction
- AnyDepth: Depth Estimation Made Easy
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- Predicting ventilation from single breathing phase non‐contrast CT using Swin Transformers
- LACC: A lightweight attention-conditional convolution network for long-term Wetland classification
- ISALux: Illumination and Segmentation Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
- OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
- Neural Enhancement of the Traditional Wang–Sheeley–Arge Solar Wind Relation
- Oral Cancer Diagnosis Using Histopathology Images: An Explainable Hybrid Transformer Framework
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- DyFormer: A Scalable Dynamic Graph Transformer with Provable Benefits on Generalization Ability
- 1st Place Solutions for UG2+ Challenge 2021 -- (Semi-)supervised Face detection in the low light condition
- Knowledge Transfer from Interaction Learning
- MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
- Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
- Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
- FROQ: Observing Face Recognition Models for Efficient Quality Assessment
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
- CSDformer: A Conversion Method for Fully Spike-Driven Transformer
- MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
- MO R-CNN: Multispectral Oriented R-CNN for Object Detection in Remote Sensing Image
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- V-CECE: Visual Counterfactual Explanations via Conceptual Edits
- FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
- Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- CGTGait: Collaborative Graph and Transformer for Gait Emotion Recognition
- Spectral Compressive Imaging via Chromaticity-Intensity Decomposition
- ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching
- Global Regulation and Excitation via Attention Tuning for Stereo Matching
- Saccadic Vision for Fine-Grained Visual Classification
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion
- CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan Reconstruction
- Region-Aware Deformable Convolutions
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- Autoguided Online Data Curation for Diffusion Model Training
- FlowCast-ODE: Continuous Hourly Weather Forecasting with Dynamic Flow Matching and ODE Solver
- HybridMamba: A Dual-domain Mamba for 3D Medical Image Segmentation
- Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation
- CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
- RaFD: Flow-Guided Radar Detection for Robust Autonomous Driving
- Google Landmark Retrieval 2021 Competition Third Place Solution
- A Multilevel Multimodal Fusion Transformer for Remote Sensing Semantic Segmentation
- TTSR: A Transformer-Based Topography Neural Network for Digital Elevation Model Super-Resolution
- Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps
- Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
- Data Leakage in Visual Datasets
- Neural Proteomics Fields for Super-resolved Spatial Proteomics Prediction
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Self Identity Mapping
- HGACNet: Hierarchical Graph Attention Network for Cross-Modal Point Cloud Completion
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation
- LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
- SAGA: Selective Adaptive Gating for Efficient and Expressive Linear Attention
- Performance is not All You Need: Sustainability Considerations for Algorithms
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- UVO Challenge on Video-based Open-World Segmentation 2021: 1st Place Solution
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- CECT-Mamba: a Hierarchical Contrast-enhanced-aware Model for Pancreatic Tumor Subtyping from Multi-phase CECT
- Aggregated Mutual Learning between CNN and Transformer for semi-supervised medical image segmentation
- Dual-stream autoencoder for channel-level multi-scale feature extraction in hyperspectral unmixing
- SPMFormer: Simplified Physical Model-based transformer with cross-space loss for underwater image enhancement
- DyGLNet: Hybrid Global-Local Feature Fusion with Dynamic Upsampling for Medical Image Segmentation
- NEFT: A Unified Transformer Framework for Efficient Near-Field CSI Feedback in XL-MIMO Systems
- Road Obstacle Video Segmentation
- Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
- ARS-DETR: Aspect Ratio-Sensitive Detection Transformer for Aerial Oriented Object Detection
- Swin Transformer-Based Multiscale Attention Model for Landslide Extraction From Large-Scale Area
- Multi Anatomy X-Ray Foundation Model
- Towards Foundational Models for Single-Chip Radar
- DS@GT AnimalCLEF: Triplet Learning over ViT Manifolds with Nearest Neighbor Classification for Animal Re-identification
- U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
- End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI
- CE-RS-SBCIT A Novel Channel Enhanced Hybrid CNN Transformer with Residual, Spatial, and Boundary-Aware Learning for Brain Tumor MRI Analysis
- RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration
- GRASP: Geospatial pixel Reasoning viA Structured Policy learning
- Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
- Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
- CCoMAML: Efficient Cattle Identification Using Cooperative Model-Agnostic Meta-Learning
- Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
- ToMA: Token Merge with Attention for Diffusion Models
- Multimodal SAM-adapter for Semantic Segmentation
- Transformer Networks for Continuous Gravitational-wave Searches
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- Compressed Video Quality Enhancement: Classifying and Benchmarking over Standards
- BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals
- Local Information Matters: A Rethink of Crowd Counting
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
- mRadNet: A Compact Radar Object Detector with MetaFormer
- Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future Trends
- Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
- XCiT: Cross-Covariance Image Transformers
- OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
- A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Live(r) Die: Predicting Survival in Colorectal Liver Metastasis
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Deep learning in multiple animal tracking: A survey
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
- First-order State Space Model for Lightweight Image Super-resolution
- CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- RepViT-CXR: A Channel Replication Strategy for Vision Transformers in Chest X-ray Tuberculosis and Pneumonia Classification
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- SEEC: Segmentation-Assisted Multi-Entropy Models for Learned Lossless Image Compression
- SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
- DR-CircuitGNN: Training Acceleration of Heterogeneous Circuit Graph Neural Network on GPUs
- H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- Hybrid Swin Attention Networks for Simultaneously Low-Dose PET and CT Denoising
- Integrated Detection and Tracking Based on Radar Range-Doppler Feature
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
- Learning spatially structured open quantum dynamics with regional-attention transformers
- MRD-LiNet: A Novel Lightweight Hybrid CNN with Gradient-Guided Unlearning for Improved Drought Stress Identification
- IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
- Florence: A New Foundation Model for Computer Vision
- AIM 2025 Challenge on High FPS Motion Deblurring: Methods and Results
- Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
- DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
- Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
- A brain-inspired paradigm for scalable quantum vision
- SpecSwin3D: Generating Hyperspectral Imagery from Multispectral Data via Transformer Networks
- MSDA-HLGCformer-based context-aware fusion network for underwater organism detection
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- HyPINO: Multi-Physics Neural Operators via HyperPINNs and the Method of Manufactured Solutions
- Accurate medium-range global weather forecasting with 3D neural networks
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- Learning A 3D-CNN and Transformer Prior for Hyperspectral Image Super-Resolution
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Efficient Video Transformers with Spatial-Temporal Token Selection
- Toward Accessible Dermatology: Skin Lesion Classification Using Deep Learning Models on Mobile-Acquired Images
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- Medical Image Segmentation Review: The Success of U-Net
- Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
- Robust End-to-End FSO Transmission with Joint Coding Modulation and BiLSTM-Based Channel Modeling under Atmospheric Turbulence
- RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
- Unsupervised Instance Segmentation with Superpixels
- LGBP-OrgaNet: Learnable Gaussian Band Pass Fusion of CNN and Transformer Features for Robust Organoid Segmentation and Tracking
- Gradient Estimation Methods of Approximate Multipliers for High-Accuracy Retraining of Deep Learning Models
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- LLMCDSR: Enhancing Cross-Domain Sequential Recommendation with Large Language Models
- RiverScope: High-Resolution River Masking Dataset
- Vision encoders should be image size agnostic and task driven
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery
- Associating Objects with Transformers for Video Object Segmentation
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Investigating Transfer Learning Capabilities of Vision Transformers and CNNs by Fine-Tuning a Single Trainable Block
- SpectMamba: Integrating Frequency and State Space Models for Enhanced Medical Image Detection
- Predicting prognosis of light-chain cardiac amyloidosis by magnetic resonance imaging and deep learning
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- DCA: Graph-Guided Deep Embedding Clustering for Brain Atlases
- Exploring Over-stationarization in Deep Learning-based Bus/Tram Arrival Time Prediction: Analysis and Non-stationary Effect Recovery
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- Satellite Image Utilization for Dehazing with Swin Transformer-Hybrid U-Net and Watershed loss
- Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition
- Encoder-Only Image Registration
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- MPANet: Motion Pattern Aggregation Network for Gait Recognition
- Efficient Diffusion Model for Image Restoration by Residual Shifting
- IA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers
- On the Integration of Self-Attention and Convolution
- FW-GAN: Frequency-Driven Handwriting Synthesis with Wave-Modulated MLP Generator
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- YOLOX: Exceeding YOLO Series in 2021
- SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- Prediction of Distant Metastasis in Head and Neck Cancer Patients Using Tumor and Peritumoral Multi-Modal Deep Learning
- Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection
- Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
- CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
- STGAtt: A Spatial-Temporal Unified Graph Attention Network for Traffic Flow Forecasting
- Foundation Models for Cross-Domain EEG Analysis Application: A Survey
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
- Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- Local Patch AutoAugment with Multi-Agent Collaboration
- Gradient Rectification for Robust Calibration under Distribution Shift
- From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Machine-learning competition to grade EEG background patterns in newborns with hypoxic-ischaemic encephalopathy
- Autoregressive Universal Video Segmentation Model
- SwinIR: Image Restoration Using Swin Transformer
- Random forest-based out-of-distribution detection for robust lung cancer segmentation
- Can we make NeRF-based visual localization privacy-preserving?
- ByteTrack: Multi-object Tracking by Associating Every Detection Box
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Huracan: A skillful end-to-end data-driven system for ensemble data assimilation and weather prediction
- A Classification Network With Coordinate and Class‐Specific Residual Attention for Accurate Detection of Heart Failure‐Related Findings in Chest X‐Rays
- VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Signals vs. Videos: Advancing Motion Intention Recognition for Human-Robot Collaboration in Construction
- D3FNet: A Differential Attention Fusion Network for Fine-Grained Road Structure Extraction in Remote Perception Systems
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
- Global Interaction Modelling in Vision Transformer via Super Tokens
- The Loupe: A Plug-and-Play Attention Module for Amplifying Discriminative Features in Vision Transformers
- You Only Pose Once: A Minimalist's Detection Transformer for Monocular RGB Category-level 9D Multi-Object Pose Estimation
- Reliable Smoke Detection via Optical Flow-Guided Feature Fusion and Transformer-Based Uncertainty Modeling
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
- MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
- TCFNet: Bidirectional face-bone transformation via a Transformer-based coarse-to-fine point movement network
- GasTwinFormer: A Hybrid Vision Transformer for Livestock Methane Emission Segmentation and Dietary Classification in Optical Gas Imaging
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- Communication-Efficient Federated Learning with Adaptive Number of Participants
- A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports
- FaPN: Feature-aligned Pyramid Network for Dense Image Prediction
- ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
- AIM 2025 challenge on Inverse Tone Mapping Report: Methods and Results
- Wavy Transformer
- Generalization vs. Memorization in Autoregressive Deep Learning: Or, Examining Temporal Decay of Gradient Coherence
- FractMorph: A Fractional Fourier-Based Multi-Domain Transformer for Deformable Image Registration
- Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems
- Geometry-Aware Video Inpainting for Joint Headset Occlusion Removal and Face Reconstruction in Social XR
- HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
- Illusions in Humans and AI: How Visual Perception Aligns and Diverges
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- MedFormer: a data-driven model for forecasting the Mediterranean Sea
- PVT v2: Improved baselines with pyramid vision transformer
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- ENA: Efficient N-dimensional Attention
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- UniNet: Unified Architecture Search with Convolution, Transformer, and MLP
- AIM: Amending Inherent Interpretability via Self-Supervised Masking
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Importance-Aware Robust Semantic Transmission for LEO Satellite-Ground Communication
- Subcortical Masks Generation in CT Images via Ensemble-Based Cross-Domain Label Transfer
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
- Fourier-Guided Attention Upsampling for Image Super-Resolution
- DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
- PSScreen: Partially Supervised Multiple Retinal Disease Screening
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Improving Learning of New Diseases through Knowledge-Enhanced Initialization for Federated Adapter Tuning
- Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Teaching LLMs to Speak Spectroscopy
- Hierarchical Graph Attention Network for No-Reference Omnidirectional Image Quality Assessment
- MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention
- NEURAL: Attention-Guided Pruning for Unified Multimodal Resource-Constrained Clinical Evaluation
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Decentralized Rank Scheduling for Energy-Constrained Multi-Task Federated Fine-Tuning in Edge-Assisted IoV Networks
- Learning Spatial Decay for Vision Transformers
- RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- AI-Driven Detection and Analysis of Handwriting on Seized Ivory: A Tool to Uncover Criminal Networks in the Illicit Wildlife Trade
- FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
- UltraLight Med-Vision Mamba for Classification of Neoplastic Progression in Tubular Adenomas
- Dual-stream Network for Visual Recognition
- Automated Charge Transition Detection in Quantum Dot Charge Stability Diagrams
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- PADReg: Physics-Aware Deformable Registration Guided by Contact Force for Ultrasound Sequences
- QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Scaling Learned Image Compression Models up to 1 Billion
- Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
- SelfHVD: Self-Supervised Handheld Video Deblurring
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
- CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Confounder Identification-free Causal Visual Feature Learning
- Scalable Swin Transformer network for brain tumor segmentation from incomplete MRI modalities
- Optimized Deep Learning for Mammography: Augmentation and Tailored Architectures
- HV-OCTAMamba: A high-order vision Mamba network for robust retinal vasculature segmentation in OCTA images
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- Text-guided Visual Prompt DINO for Generic Segmentation
- UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
- Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
- Hybrid(Transformer+CNN)-based Polyp Segmentation
- Lightweight Quad Bayer HybridEVS Demosaicing via State Space Augmented Cross-Attention
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- SMOL-MapSeg: Show Me One Label as prompt
- Deformable Attention Graph Representation Learning for Histopathology Whole Slide Image Analysis
- CT-GRAPH: Hierarchical Graph Attention Network for Anatomy-Guided CT Report Generation
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
- Contextual Object Detection with Multimodal Large Language Models
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Unified modality separation: A vision-language framework for unsupervised domain adaptation
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image Compression
- Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
- Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
- Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- Benchmarking Foundation Models for Mitotic Figure Classification
- Efficient Inter-Task Attention for Multitask Transformer Models
- VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Deeper Inside Deep ViT
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
- Towards Globally Predictable k-Space Interpolation: A White-box Transformer Approach
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
- AttZoom: Attention Zoom for Better Visual Features
- Hidden Dynamics of Massive Activations in Transformer Training
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation
- Monocular Depth Estimation with Global-Aware Discretization and Local Context Modeling
- SSFMamba: Symmetry-driven Spatial-Frequency Feature Fusion for 3D Medical Image Segmentation
- CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
- Rethinking Transparent Object Grasping: Depth Completion with Monocular Depth Estimation and Instance Mask
- TRUDI and TITUS: A Multi-Perspective Dataset and A Three-Stage Recognition System for Transportation Unit Identification
- DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal
- S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
- Context Guided Transformer Entropy Modeling for Video Compression
- Large Kernel MedNeXt for Breast Tumor Segmentation and Self-Normalizing Network for pCR Classification in Magnetic Resonance Images
- DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- CGCCE-Net:Change-Guided Cross Correlation Enhancement Network for Remote Sensing Building Change Detection
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Self-Navigated Residual Mamba for Universal Industrial Anomaly Detection
- LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Skip priors and add graph-based anatomical information, for point-based Couinaud segmentation
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network
- SWAN: Synergistic Wavelet-Attention Network for Infrared Small Target Detection
- RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification
- Conquering High Packet-Loss Erasure: MoE Swin Transformer-Based Video Semantic Communication
- Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- A Framework Combining 3D CNN and Transformer for Video-Based Behavior Recognition
- Flow Matching for Probabilistic Learning of Dynamical Systems from Missing or Noisy Data
- OmniUnet: A Multimodal Network for Unstructured Terrain Segmentation on Planetary Rovers Using RGB, Depth, and Thermal Imagery
- DBLP: Noise Bridge Consistency Distillation For Efficient And Reliable Adversarial Purification
- Representation Shift: Unifying Token Compression with FlashAttention
- VQ-DeepISC: Vector Quantized-Enabled Digital Semantic Communication with Channel Adaptive Image Transmission
- Multimodal Referring Segmentation: A Survey
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space Model
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Adjustable Spatio-Spectral Hyperspectral Image Compression Network
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- Are Transformers Effective for Time Series Forecasting?
- Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors
- DACA-Net: A Degradation-Aware Conditional Diffusion Network for Underwater Image Enhancement
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
- Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- Whole-brain Transferable Representations from Large-Scale fMRI Data Improve Task-Evoked Brain Activity Decoding
- MSQ: Memory-Efficient Bit Sparsification Quantization
- Learning from Heterogeneous Structural MRI via Collaborative Domain Adaptation for Late-Life Depression Assessment
- Vision-Language Cross-Attention for Real-Time Autonomous Driving
- From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
- Spatial-Temporal-Spectral Mamba with Sparse Deformable Token Sequence for Enhanced MODIS Time Series Classification
- Brain Tumor Segmentation in Sub-Sahara Africa with Advanced Transformer and ConvNet Methods: Fine-Tuning, Data Mixing and Ensembling
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Color as the Impetus: Transforming Few-Shot Learner
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning
- Staining and locking computer vision models without retraining
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
- SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
- Cross-Architecture Distillation Made Simple with Redundancy Suppression
- Unlocking Interpretability for RF Sensing: A Complex-Valued White-Box Transformer
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation
Related