Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
2021/03/25 by Ze Liu, Yutong Lin, Liu, Ze +13 · 2 voices · 2,132 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Computation #Computer science #Computer vision #Domain Adaptation and Few-Shot Learning #Electrical engineering #Engineering #Segmentation #Transformer #Voltage #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2103.14030
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2021/03/25 · arxiv created 2021/08/17 · arxiv updated 2021/08/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at~\urlhttps://github.com/microsoft/Swin-Transformer.
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
- Unsupervised Multimodal Intent Discovery via MLLM-Guided Concept Generation and Semantic Propagation
- Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
- Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation
- ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors
- CARE: Anti-entanglement Ultrasound Image Segmentation via Channel-Aware Region Extrication
- IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing
- Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
- CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras
- Spatial Semantic Communication: When Semantic Transmission Meets Index Modulation
- Conjugate Gradient Unrolled Network with PSF Conditioning for Non-Diagonal Data Fidelity in CASSI Reconstruction
- Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
- AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries
- G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
- MedDiT4SR: Tri-Stream Joint Adaptation of Pre-Trained Diffusion Transformers for Medical Image Super-Resolution
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
- Medical Imaging Fusing Vision Transformer: Laryngeal Cancer Screening with Explanation
- Scaling Synthetic-Image Pre-Training for Federated Fine-Tuning of Large Vision Models
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
- MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
- Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention
- Open-Vocabulary Gaze Object Prediction: Benchmark and Method
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection
- SkyEV: RGB-Event UAV detection and tracking dataset and baseline
- Privileged Lesion-Context Relational Distillation for Mask-Free Skin Lesion Classification
- InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation
- DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
- Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection
- BMFA: Boundary-Minority Free-Energy Adaptive Screening
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Tumor-anchored deep feature random forests for out-of-distribution detection in lung cancer segmentation
- BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
- BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
- High-Capacity Robust Watermarking Technology for High-Resolution Images
- Overlapping Schwarz Attention: Hierarchical Attention via Domain Decomposition
- Constraint-Anchored Reasoning Traces
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
- Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery
- DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
- An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
- VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
- Avoiding Dilution: Using Diffusion and Vision Transformers to resolve Majorana Features in Nanowires at High Temperature
- The Third Competition on Document Forgery Detection on ID-Cards and Passports
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
- Still image and spatial-temporal tomato data enabling detection, segmentation, tracking, and video-instance segmentation using strong and weak labels
- Angular Gaussian Supervised Contrastive Learning for Long-Tailed Electrocardiogram Arrhythmia Diagnosis
- High-Fidelity Mural Restoration via a Unified Hybrid Mask-Aware Transformer
- The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers
- Towards Hierarchical Structure Understanding of Newspaper Images
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
- SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
- Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- FISHER: Gradient-Decoupled Hierarchical Multi-Task Learning for Fine-Grained Aquatic Species Recognition
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- A Survey on GNN ‐Based Link Prediction: Techniques, Applications, and Challenges
- Federated Deep Subspace Clustering
- MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
- Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
- BreastDCEDL: A standardized deep learning-ready breast DCE-MRI dataset of 2,070 patients
- A vision-language conditioned physics-aware imitation learning approach for bimanual robotic dexterous assembly
- Spectral and Spatial Graph Learning for Multispectral Solar Image Compression
- The Mean-Field Dynamics of Transformers
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- Cameras as Relative Positional Encoding
- JAFAR: Jack up Any Feature at Any Resolution
- FLAIR-HUB: Large-scale Multimodal Dataset for Land Cover and Crop Mapping
- NeXtBrain: Combining local and global feature learning for brain tumor classification
- Leaner Transformers: More Heads, Less Depth
- Fourier-mixed window attention for efficient and robust long sequence time-series forecasting
- DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation
- Perception Encoder: The best visual embeddings are not at the output of the network
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
- Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical Systems
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
- Multi-label Classification with Panoptic Context Aggregation Networks
- BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
- CountGD++: Generalized Prompting for Open-World Counting
- ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- SwinCCIR: An end-to-end deep network for Compton camera imaging reconstruction
- Revisiting [CLS] and Patch Token Interaction in Vision Transformers
- Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- SCAFusion: A Multimodal 3D Detection Framework for Small Object Detection in Lunar Surface Exploration
- Bright 4B: Scaling Hyperspherical Learning for Segmentation in 3D Brightfield Microscopy
- Feature Learning with Multi-Stage Vision Transformers on Inter-Modality HER2 Status Scoring and Tumor Classification on Whole Slides
- SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception
- DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
- Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments
- AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
- A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting
- Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications
- Neuromorphic Object Detection: An In-Depth Study and Future Directions
- Multi‐angle, cross‐domain fusion strategy enhances automated insect identification and hierarchical categorization: a case study on assassin bugs (Hemiptera: Reduviidae)
- Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
- RhythmFormer: Extracting Patterned rPPG Signals based on Periodic Sparse Attention
- SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- Physics Transformer: Tailoring Transformer for General PDE Prediction
- DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
- WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
- ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
- Optimization of Collaborative Semantic Communication Network Performance with Channel and Content Preference Feedback
- Agentic Autoresearch for CT Reconstruction
- Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention
- Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy
- Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
- FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog
- MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion
- DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification
- DDSNet: Dual-domain Symmetry-aware Network for PCSEL Property Prediction
- UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation
- Patch-Discontinuity Mining for Generalized Deepfake Detection
- Benchmarking deep learning models for Raman spectroscopy across open-source datasets
- A Lightweight Multi-Scale Attention Framework for Real-Time Spinal Endoscopic Instance Segmentation
- Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
- SLIM-Brain: A Data- and Training-Efficient Foundation Model for fMRI Data Analysis
- SDUM: A Scalable Deep Unrolled Model for Universal MRI Reconstruction
- Vision Transformers are Circulant Attention Learners
- RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
- Fast SAM2 with Text-Driven Token Pruning
- Granular Ball Guided Masking: Structure-aware Data Augmentation
- Knowledge-Driven 3D Semantic Spectrum Map: KE-VQ-Transformer Based UAV Semantic Communication and Map Completion
- X-ray Insights Unleashed: Pioneering the Enhancement of Multi-Label Long-Tail Data
- DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction
- Simplifying Multi-Task Architectures Through Task-Specific Normalization
- USE: A Unified Model for Universal Sound Separation and Extraction
- Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks
- Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Field-Space Attention for Structure-Preserving Earth System Transformers
- Degradation-Aware Metric Prompting for Hyperspectral Image Restoration
- HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
- SemanticGen: Video Generation in Semantic Space
- Vehicle-centric Perception via Multimodal Structured Pre-training
- Non-Contrast CT Esophageal Varices Grading through Clinical Prior-Enhanced Multi-Organ Analysis
- On Network-Aware Semantic Communication and Edge-Cloud Collaborative Intelligence Systems
- Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration
- Mamba-Based Modality Disentanglement Network for Multi-Contrast MRI Reconstruction
- Efficient Vision Mamba for MRI Super-Resolution via Hybrid Selective Scanning
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- HyGE-Occ: Hybrid View-Transformation with 3D Gaussian and Edge Priors for 3D Panoptic Occupancy Prediction
- Overcoming Spectral Bias via Cross-Attention
- PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
- Object-Centric Framework for Video Moment Retrieval
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- A two-stream network with global-local feature fusion for bone age assessment
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- WDFFU-Mamba: A Wavelet-guided Dual-attention Feature Fusion Mamba for Breast Tumor Segmentation in Ultrasound Images
- Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts
- SARMAE: Masked Autoencoder for SAR Representation Learning
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- LAPX: Lightweight Hourglass Network with Global Context
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
- SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal
- Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
- Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
- Enhancing Visual Sentiment Analysis via Semiotic Isotopy-Guided Dataset Construction
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
- SS4D: Native 4D Generative Model via Structured Spacetime Latents
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- rNCA: Self-Repairing Segmentation Masks
- USTM: Unified Spatial and Temporal Modeling for Continuous Sign Language Recognition
- Anatomy Guided Coronary Artery Segmentation from CCTA Using Spatial Frequency Joint Modeling
- Referring Change Detection in Remote Sensing Imagery
- Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
- Multi-temporal Calving Front Segmentation
- Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Stronger Normalization-Free Transformers
- Graph Laplacian Transformer with Progressive Sampling for Prostate Cancer Grading
- Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
- Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment
- Diffusion Is Your Friend in Show, Suggest and Tell
- Towards Visual Re-Identification of Fish using Fine-Grained Classification for Electronic Monitoring in Fisheries
- CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
- A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- FBA2D: Frequency-based Black-box Attack for AI-generated Image Detection
- DistillFSS: Synthesizing Few-Shot Knowledge into a Lightweight Segmentation Model
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- Skewness-Guided Pruning of Multimodal Swin Transformers for Federated Skin Lesion Classification on Edge Devices
- C-DIRA: Computationally Efficient Dynamic ROI Routing and Domain-Invariant Adversarial Learning for Lightweight Driver Behavior Recognition
- HATSolver: Learning Groebner Bases with Hierarchical Attention Transformers
- Residual-SwinCA-Net: A Channel-Aware Integrated Residual CNN-Swin Transformer for Malignant Lesion Segmentation in BUSI
- Fast-BEV++: Fast by Algorithm, Deployable by Design
- SOP2: Transfer Learning with Scene-Oriented Prompt Pool on 3D Object Detection
- Fourier-RWKV: A Multi-State Perception Network for Efficient Image Dehazing
- Accelerated Rotation-Invariant Convolution for UAV Image Segmentation
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Identification of Deforestation Areas in the Amazon Rainforest Using Change Detection Models
- Generalized Referring Expression Segmentation on Aerial Photos
- DGGAN: Degradation Guided Generative Adversarial Network for Real-time Endoscopic Video Enhancement
- Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- Omni-Referring Image Segmentation
- Rectifying Latent Space for Generative Single-Image Reflection Removal
- Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
- When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image Segmentation
- 3D Path Planning for Robot-assisted Vertebroplasty from Arbitrary Bi-plane X-ray via Differentiable Rendering
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- Image Semantic Communication with Quadtree Partition-based Coding
- Self-Supervised Learning for Transparent Object Depth Completion Using Depth from Non-Transparent Objects
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- Multi Task Denoiser Training for Solving Linear Inverse Problems
- Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- Shift-Window Meets Dual Attention: A Multi-Model Architecture for Specular Highlight Removal
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
- On the Temporality for Sketch Representation Learning
- Dual Cross-Attention Siamese Transformer for Rectal Tumor Regrowth Assessment in Watch-and-Wait Endoscopy
- MKSNet: Advanced Small Object Detection in Remote Sensing Imagery with Multi-Kernel and Dual Attention Mechanisms
- HBFormer: A Hybrid-Bridge Transformer for Microtumor and Miniature Organ Segmentation
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching
- Hierarchical Attention for Sparse Volumetric Anomaly Detection in Subclinical Keratoconus
- When and how to automate image analysis for wildlife monitoring? Guidelines and lessons from a worked example of seabirds in a dynamic coastal environment
- DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
- GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
- Layout Anything: One Transformer for Universal Room Layout Estimation
- Defense That Attacks: How Robust Models Become Better Attackers
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
- Data-Centric Visual Development for Self-Driving Labs
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- Robust Rigid and Non-Rigid Medical Image Registration Using Learnable Edge Kernels
- On the Unreasonable Effectiveness of Last-layer Retraining
- ViT3: Unlocking Test-Time Training in Vision
- Toward Content-based Indexing and Retrieval of Head and Neck CT with Abscess Segmentation
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
- nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
- SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
- LAHNet: Local Attentive Hashing Network for Point Cloud Registration
- Joint Multi-scale Gated Transformer and Prior-guided Convolutional Network for Learned Image Compression
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- Cross-Domain Federated Semantic Communication with Global Representation Alignment and Domain-Aware Aggregation
- Silhouette-based Gait Foundation Model
- Structured Context Learning for Generic Event Boundary Detection
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
- Deep Learning for Restoring MPI System Matrices Using Simulated Training Data
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla
- Stable-Drift: A Patient-Aware Latent Drift Replay Method for Stabilizing Representations in Continual Learning
- GSPN-2: Efficient Parallel Sequence Modeling
- Learning to Predict Aboveground Biomass from RGB Images with 3D Synthetic Scenes
- Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Hard Spatial Gating for Precision-Driven Brain Metastasis Segmentation: Addressing the Over-Segmentation Paradox in Deep Attention Networks
- IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Small Object Detection for Birds with Swin Transformer
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Lost in Time? A Meta-Learning Framework for Time-Shift-Tolerant Physiological Signal Transformation
- Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
- PathMamba: A Hybrid Mamba-Transformer for Topologically Coherent Road Segmentation in Satellite Imagery
- Open Vocabulary Compositional Explanations for Neuron Alignment
- Intriguing Properties of Dynamic Sampling Networks
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics
- LiMT: A Multi-task Liver Image Benchmark Dataset
- HybriDLA: Hybrid Generation for Document Layout Analysis
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Neural surrogates for designing gravitational wave detectors
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
- DEAP-3DSAM: Decoder Enhanced and Auto Prompt SAM for 3D Medical Image Segmentation
- Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
- Changes in Gaza: DINOv3-Powered Multi-Class Change Detection for Damage Assessment in Conflict Zones
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
- Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
- Robust Nonlinear Transform Coding: A Framework for Generalizable Joint Source-Channel Coding
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- From Features to Reference Points: Lightweight and Adaptive Fusion for Cooperative Autonomous Driving
- OceanForecastBench: A Benchmark Dataset for Data-Driven Global Ocean Forecasting
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
- Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
- LRDUN: A Low-Rank Deep Unfolding Network for Efficient Spectral Compressive Imaging
- PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Compact neural networks for astronomy with optimal transport bias correction
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- TransLK-Net: Entangling Transformer and Large Kernel for Progressive and Collaborative Feature Encoding and Decoding in Medical Image Segmentation
- SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
- Feature Partitioning and Semantic Equalization for Intrinsic Robustness in Semantic Communication under Packet Loss
- Fine-grained MoE Load Balancing with Linear Programming
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- A Spatial Semantics and Continuity Perception Attention for Remote Sensing Water Body Change Detection
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
- CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking
- A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
- WarNav: An Autonomous Driving Benchmark for Segmentation of Navigable Zones in War Scenes
- Communication-Pipelined Split Federated Learning for Foundation Model Fine-Tuning in UAV Networks
- IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging
- FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
- Parameter Aware Mamba Model for Multi-task Dense Prediction
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Segment Anything Across Shots: A Method and Benchmark
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
- CD-DPE: Dual-Prompt Expert Network Based on Convolutional Dictionary Feature Decoupling for Multi-Contrast MRI Super-Resolution
- Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- Semi-Supervised Multi-Task Learning for Interpretable Quality As- sessment of Fundus Images
- MRIQT: Physics-Aware Diffusion Model for Image Quality Transfer in Neonatal Ultra-Low-Field MRI
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
- HDW-SR: High-Frequency Guided Diffusion Model based on Wavelet Decomposition for Image Super-Resolution
- CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
- DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
- H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
- SAGE: Saliency-Guided Contrastive Embeddings
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- Counting Through Occlusion: Framework for Open World Amodal Counting
- SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
- MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
- MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging
- DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
- SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
- Application of Graph Based Vision Transformers Architectures for Accurate Temperature Prediction in Fiber Specklegram Sensors
- Heterogeneous Complementary Distillation
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
- Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
- LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
- DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
- MPCM-Net: Multi-scale network integrates partial attention convolution with Mamba for ground-based cloud image segmentation
- Stratified Knowledge-Density Super-Network for Scalable Vision Transformers
- Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
- FAST-CAD: A Fairness-Aware Framework for Non-Contact Stroke Diagnosis
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- SSMRadNet : A Sample-wise State-Space Framework for Efficient and Ultra-Light Radar Segmentation and Object Detection
- How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI
- MVSMamba: Multi-View Stereo with State Space Model
- Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
- Positive Semi-definite Latent Factor Grouping-Boosted Cluster-reasoning Instance Disentangled Learning for WSI Representation
- REASON: Probability map-guided dual-branch fusion framework for gastric content assessment
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Rethinking Explanation Evaluation under the Retraining Scheme
- Pixel-level Quality Assessment for Oriented Object Detection
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving
- CSF-Net: Context-Semantic Fusion Network for Large Mask Inpainting
- The Impact of Longitudinal Mammogram Alignment on Breast Cancer Risk Assessment
- Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Detecting Generated Images by Fitting Natural Image Distributions
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
- CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
- BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
- LeCoT: revisiting network architecture for two-view correspondence pruning
- CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
- Anatomy-Aware Lymphoma Lesion Detection in Whole-Body PET/CT
- Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
- MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
- REOcc: Camera-Radar Fusion with Radar Feature Enrichment for 3D Occupancy Prediction
- Active Learning for Animal Re-Identification with Ambiguity-Aware Sampling
- Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation
- Real-Time LiDAR Super-Resolution via Frequency-Aware Multi-Scale Fusion
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- On Modality Incomplete Infrared-Visible Object Detection: An Architecture Compatibility Perspective
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
- LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation
- Physics-Informed Image Restoration via Progressive PDE Integration
- MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
- Temporal-Guided Visual Foundation Models for Event-Based Vision
- ReMoD: Rethinking Modality Contribution in Multimodal Stance Detection via Dual Reasoning
- CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
- HarmoQ: Harmonized Post-Training Quantization for High-Fidelity Image
- MACMD: Multi-dilated Contextual Attention and Channel Mixer Decoding for Medical Image Segmentation
- Hilbert-Guided Sparse Local Attention
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- Rethinking Metrics and Diffusion Architecture for 3D Point Cloud Generation
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- SurgiATM: A Physics-Guided Plug-and-Play Model for Deep Learning-Based Smoke Removal in Laparoscopic Surgery
- Learning Fourier shapes to probe the geometric world of deep neural networks
- What's on Your Plate? Inferring Chinese Cuisine Intake from Wearable IMUs
- No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
- Global 3D Reconstruction of Clouds & Tropical Cyclones
- Landslide Hazard Mapping with Geospatial Foundation Models: Geographical Generalizability, Data Scarcity, and Band Adaptability
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- Evaluating the Impact of Weather-Induced Sensor Occlusion on BEVFusion for 3D Object Detection
- Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
- When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation
- Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
- Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge Computing
- Image-Intrinsic Priors for Integrated Circuit Defect Detection and Novel Class Discovery via Self-Supervised Learning
- Individual identification of brown bears using pose-aware metric learning
- Generative Hints
- Domain-Adaptive Transformer for Data-Efficient Glioma Segmentation in Sub-Saharan MRI
- UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- GAFD-CC: Global-Aware Feature Decoupling with Confidence Calibration for OOD Detection
- Fast Measuring Pavement Crack Width by Cascading Principal Component Analysis
- Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Towards Automated Petrography
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- Region-Aware Reconstruction Strategy for Pre-training fMRI Foundation Model
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
- FPS: Feedforward-based Parameter Selection For Efficient Fine-Tuning
- AFM-Net: Advanced Fusing Hierarchical CNN Visual Priors with Global Sequence Modeling for Remote Sensing Image Scene Classification
- Hierarchical Transformers for Unsupervised 3D Shape Abstraction
- FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
- Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
- SYNAPSE-Net: A Unified Framework with Lesion-Aware Hierarchical Gating for Robust Segmentation of Heterogeneous Brain Lesions
- PF-DAformer: Proximal Femur Segmentation via Domain Adaptive Transformer for Dual-Center QCT
- SA2Net: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging
- SPG-CDENet: Spatial Prior-Guided Cross Dual Encoder Network for Multi-Organ Segmentation
- Multi-hop Parallel Image Semantic Communication for Distortion Accumulation Mitigation
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
- Regime identification and control of extremes in the non-autonomous Lorenz model with chaos and intransitivity
- Synthetic Data Reveals Generalization Gaps in Correlated Multiple Instance Learning
- Leveraging an Atmospheric Foundational Model for Subregional Sea Surface Temperature Forecasting
- SPADE: Sparsity Adaptive Depth Estimator for Zero-Shot, Real-Time, Monocular Depth Estimation in Underwater Environments
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- Self-distillation with Batch Knowledge Ensembling Improves ImageNet Classification
- A Comparison of Data Augmentation Methods for Training Deep Neural Networks on Synthetic Aperture Sonar
- CASIAL: Geometric Distortion Robust Image Watermarking
- Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Bridging Global Context Interactions for High-Fidelity Image Completion
- BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens
- An Image Patch is a Wave: Phase-Aware Vision MLP
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- DCIRNet: Depth Completion with Iterative Refinement for Dexterous Grasping of Transparent and Reflective Objects
- ViT-Transformer: Self-attention mechanism based constitutive modeling for nonlinear heterogeneous materials
- Attentive multilayer fusion for vision transformers
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- A multimodal whole-slide foundation model for pathology
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Energy-Efficient Autonomous Driving with Adaptive Perception and Robust Decision
- Test-Time Adaptive Object Detection with Foundation Model
- Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- DRIP: Dynamic patch Reduction via Interpretable Pooling
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Decoupling What to Count and Where to See for Referring Expression Counting
- Unlocking Out-of-Distribution Generalization in Dynamics through Physics-Guided Augmentation
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- UniField: Joint Multi-Domain Training for Universal Surface Pressure Modeling
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- Revealing the Potential of Learnable Perturbation Ensemble Forecast Model for Tropical Cyclone Prediction
- A Survey on Efficient Vision-Language-Action Models
- Provable test-time adaptivity and distributional robustness of in-context learning
- Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Transforming volcanic monitoring: A dataset and benchmark for onboard volcano activity detection
- Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- DAMap: Distance-aware MapNet for High Quality HD Map Construction
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- PSScreen V2: Partially Supervised Multiple Retinal Disease Screening
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- LO-SDA: Latent Optimization for Score-based Atmospheric Data Assimilation
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- Efficient Large-Deformation Medical Image Registration via Recurrent Dynamic Correlation
- Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
- Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You Need
- Simplifying Knowledge Transfer in Pretrained Models
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- FrameShield: Adversarially Robust Video Anomaly Detection
- AutoOpt: A Dataset and a Unified Framework for Automating Optimization Problem Solving
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Relieving the Over-Aggregating Effect in Graph Transformers
- LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
- Memory Constrained Dynamic Subnetwork Update for Transfer Learning
- Focal Modulation and Bidirectional Feature Fusion Network for Medical Image Segmentation
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- SutureBot: A Precision Framework & Benchmark For Autonomous End-to-End Suturing
- Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
- FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking
- Guiding diffusion models to reconstruct flow fields from sparse data
- Study of Training Dynamics for Memory-Constrained Fine-Tuning
- DARE: A Deformable Adaptive Regularization Estimator for Learning-Based Medical Image Registration
- Seabed-Net: A multi-task network for joint bathymetry estimation and seabed classification from remote sensing imagery in shallow waters
- SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
- AegisRF: Adversarial Perturbations Guided with Sensitivity for Protecting Intellectual Property of Neural Radiance Fields
- Matrix-Free Least Squares Solvers: Values, Gradients, and What to Do With Them
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- Detection and Simulation of Urban Heat Islands Using a Fine-Tuned Geospatial Foundation Model for Microclimate Impact Prediction
- A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- ProLAP: Probabilistic Language-Audio Pre-Training
- MCANet: A Coherent Multimodal Collaborative Attention Network for Advanced Modulation Recognition in Adverse Noisy Environments
- Integrated representational signatures strengthen specificity in brains and models
- Δt-Mamba3D: A Time-Aware Spatio-Temporal State-Space Model for Breast Cancer Risk Prediction
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- Rethinking PCA Through Duality
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Facial Expression-based Parkinson's Disease Severity Diagnosis via Feature Fusion and Adaptive Class Balancing
- M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- FER-YOLO: a YOLO-based classifier for facial expression recognition
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- Confidence-Weighted Semi-Supervised Learning for Skin Lesion Segmentation Using Hybrid CNN-Transformer Networks
- BARL: Bilateral Alignment in Representation and Label Spaces for Semi-Supervised Volumetric Medical Image Segmentation
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
- Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Symmetric Entropy-Constrained Video Coding for Machines
- Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
- UKANFormer: Noise-Robust Semantic Segmentation for Coral Reef Mapping via a Kolmogorov-Arnold Network-Transformer Hybrid
- CARDIUM: Congenital Anomaly Recognition with Diagnostic Images and Unified Medical records
- Attention-guided few-shot learning for metal surface defect classification
- Cost Savings from Automatic Quality Assessment of Generated Images
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- EuroMineNet: A Multitemporal Sentinel-2 Benchmark for Spatiotemporal Mining Footprint Analysis in the European Union (2015-2024)
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
- DRBD-Mamba for Robust and Efficient Brain Tumor Segmentation with Analytical Insights
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- LOTA: Bit-Planes Guided AI-Generated Image Detection
- PU-Transformer: Point Cloud Upsampling Transformer
- Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
- MetaFormer Is Actually What You Need for Vision
- MAPS: Masked Attribution-based Probing of Strategies- A computational framework to align human and model explanations
- Chimera: State Space Models Beyond Sequences
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network
- Exploring and Leveraging Class Vectors for Classifier Editing
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- Source-Free Object Detection with Detection Transformer
- MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
- Learning Model Representations Using Publicly Available Model Hubs
- VGDM: Vision-Guided Diffusion Model for Brain Tumor Detection and Segmentation
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks
- TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading
- SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
- Holistic Order Prediction in Natural Scenes
- 3D Reconstruction from Transient Measurements with Time-Resolved Transformer
- Efficient Resource-Constrained Training of Transformers via Subspace Optimization
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
- Vision Language Models: A Survey of 26K Papers
- VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
- SatFusion: A Unified Framework for Enhancing Remote Sensing Images via Multi-Frame and Multi-Source Images Fusion
- Robust Canonicalization through Bootstrapped Data Re-Alignment
- SkipSR: Faster Super Resolution with Token Skipping
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- NNDM: NNUNet Diffusion Model for Brain Tumor Segmentation
- Knowledge-Aware Mamba for Joint Change Detection and Classification from MODIS Times Series
- Pruning Self-attentions into Convolutional Layers in Single Path
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Automated Neural Architecture Design for Industrial Defect Detection
- Spatial Uncertainty Quantification in Wildfire Forecasting for Climate-Resilient Emergency Planning
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
- HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
- Lung Infection Severity Prediction Using Transformers with Conditional TransMix Augmentation and Cross-Attention
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- TransFIRA: Transfer Learning for Face Image Recognizability Assessment
- Zeeman: A Deep Learning Regional Atmospheric Chemistry Transport Model
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Critical attention scaling in long-context transformers
- Human Action Recognition from Point Clouds over Time
- A Total Variation Regularized Framework for Epilepsy-Related MRI Image Segmentation
- Diffusion2: Turning 3D Environments into Radio Frequency Heatmaps
- T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet Extraction
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- TinyViT-Batten: Few-Shot Vision Transformer with Explainable Attention for Early Batten-Disease Detection on Pediatric MRI
- Detection of retinal diseases using an accelerated reused convolutional network
- Learning more physically realistic dynamics in machine-learning based weather forecasting with latent-space constraints
- ReTiDe: Real-Time Denoising for Energy-Efficient Motion Picture Processing with FPGAs
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
- MambaCAFU: Hybrid Multi-Scale and Multi-Attention Model with Mamba-Based Fusion for Medical Image Segmentation
- Allocation of Parameters in Transformers
- Referring Expression Comprehension for Small Objects
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- CVSM: Contrastive Vocal Similarity Modeling
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Image Generation Based on Image Style Extraction
- TextCAM: Explaining Class Activation Map with Text
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- Oriented RepPoints for Aerial Object Detection
- LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
- Remote Auditing: Design-based Tests of Randomization, Selection, and Missingness with Broadly Accessible Satellite Imagery
- TransMix: Attend to Mix for Vision Transformers
- Data driven approaches in nanophotonics: A review of AI-enabled metadevices
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- Transformer Classification of Breast Lesions: The BreastDCEDLAMBL Benchmark Dataset and 0.92 AUC Baseline
- PRISM: Progressive Rain removal with Integrated State-space Modeling
- AttriGen: Automated Multi-Attribute Annotation for Blood Cell Datasets
- Indirect Attention: Turning Context Misalignment into a Feature
- VRWKV-Editor: Reducing quadratic complexity in transformer-based video editing
- The Impact of Scaling Training Data on Adversarial Robustness
- ProbMed: A Probabilistic Framework for Medical Multimodal Binding
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Bayesian Transformer for Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
- LayerD: Decomposing Raster Graphic Designs into Layers
- Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
- BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- DRCP: Diffusion on Reinforced Cooperative Perception for Perceiving Beyond Limits
- Accurate Cobb Angle Estimation via SVD-Based Curve Detection and Vertebral Wedging Quantification
- DRIFT-Net: A Spectral--Coupled Neural Operator for PDEs Learning
- Mask Clustering-based Annotation Engine for Large-Scale Submeter Land Cover Mapping
- An Enhanced Pyramid Feature Network Based on Long-Range Dependencies for Multi-Organ Medical Image Segmentation
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- FSDENet: A Frequency and Spatial Domains based Detail Enhancement Network for Remote Sensing Semantic Segmentation
- BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
- An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
- Variable Rate Image Compression via N-Gram Context based Swin-transformer
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
- Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Enhanced Fracture Diagnosis Based on Critical Regional and Scale Aware in YOLO
- Graph Your Own Prompt
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Epsilon-Scheduling
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer
- Seeing Isn't Believing: Context-Aware Adversarial Patch Synthesis via Conditional GAN
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- LizardLens: A Two-Stage Deep Learning Pipeline for Detecting and Classifying Similar Species in Visually Complex Environments
- Orochi: Versatile Biomedical Image Processor
- CCNeXt: An Effective Self-Supervised Stereo Depth Estimation Approach
- Category Discovery: An Open-World Perspective
- Integrating Background Knowledge in Medical Semantic Segmentation with Logic Tensor Networks
- Deep Learning-Based Cross-Anatomy CT Synthesis Using Adapted nnResU-Net with Anatomical Feature Prioritized Loss
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining
- Motion-Aware Transformer for Multi-Object Tracking
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
- Punching Above Precision: Small Quantized Model Distillation with Learnable Regularizer
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- WDformer: A Wavelet-based Differential Transformer Model for Time Series Forecasting
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning
- Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
- HiPerformer: A High-Performance Global-Local Segmentation Model with Modular Hierarchical Fusion Strategy
- Downscaling climate projections to 1 km with single-image super resolution
- Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- Timeliness-Aware Joint Source and Channel Coding for Adaptive Image Transmission
- Towards Self-Supervised Foundation Models for Critical Care Time Series
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Myosotis: structured computation for attention like layer
- Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
- CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography
- Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
- A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- AnyDepth: Depth Estimation Made Easy
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- ISALux: Illumination and Segmentation Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
- Predicting ventilation from single breathing phase non‐contrast CT using Swin Transformers
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- Oral Cancer Diagnosis Using Histopathology Images: An Explainable Hybrid Transformer Framework
- SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- Knowledge Transfer from Interaction Learning
- MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
- Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
- Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
- FROQ: Observing Face Recognition Models for Efficient Quality Assessment
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
- CSDformer: A Conversion Method for Fully Spike-Driven Transformer
- MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
- MO R-CNN: Multispectral Oriented R-CNN for Object Detection in Remote Sensing Image
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- V-CECE: Visual Counterfactual Explanations via Conceptual Edits
- FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
- Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- CGTGait: Collaborative Graph and Transformer for Gait Emotion Recognition
- Spectral Compressive Imaging via Chromaticity-Intensity Decomposition
- ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching
- Global Regulation and Excitation via Attention Tuning for Stereo Matching
- Opportunities and Challenges in Applying AI to Evolutionary Morphology
- Saccadic Vision for Fine-Grained Visual Classification
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion
- CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan Reconstruction
- Region-Aware Deformable Convolutions
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
- Autoguided Online Data Curation for Diffusion Model Training
- FlowCast-ODE: Continuous Hourly Weather Forecasting with Dynamic Flow Matching and ODE Solver
- HybridMamba: A Dual-domain Mamba for 3D Medical Image Segmentation
- Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation
- CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
- RaFD: Flow-Guided Radar Detection for Robust Autonomous Driving
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps
- Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
- Data Leakage in Visual Datasets
- Neural Proteomics Fields for Super-resolved Spatial Proteomics Prediction
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Self Identity Mapping
- HGACNet: Hierarchical Graph Attention Network for Cross-Modal Point Cloud Completion
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation
- Intriguing Properties of Vision Transformers
- SAGA: Selective Adaptive Gating for Efficient and Expressive Linear Attention
- Performance is not All You Need: Sustainability Considerations for Algorithms
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- CECT-Mamba: a Hierarchical Contrast-enhanced-aware Model for Pancreatic Tumor Subtyping from Multi-phase CECT
- PointMixer: MLP-Mixer for Point Cloud Understanding
- PatchFormer: An Efficient Point Transformer with Patch Attention
- DyGLNet: Hybrid Global-Local Feature Fusion with Dynamic Upsampling for Medical Image Segmentation
- NEFT: A Unified Transformer Framework for Efficient Near-Field CSI Feedback in XL-MIMO Systems
- Road Obstacle Video Segmentation
- Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
- Multi Anatomy X-Ray Foundation Model
- Towards Foundational Models for Single-Chip Radar
- DS@GT AnimalCLEF: Triplet Learning over ViT Manifolds with Nearest Neighbor Classification for Animal Re-identification
- U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
- End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI
- CE-RS-SBCIT A Novel Channel Enhanced Hybrid CNN Transformer with Residual, Spatial, and Boundary-Aware Learning for Brain Tumor MRI Analysis
- RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration
- GRASP: Geospatial pixel Reasoning viA Structured Policy learning
- Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
- Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
- CCoMAML: Efficient Cattle Identification Using Cooperative Model-Agnostic Meta-Learning
- Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
- ToMA: Token Merge with Attention for Diffusion Models
- Multimodal SAM-adapter for Semantic Segmentation
- Transformer Networks for Continuous Gravitational-wave Searches
- DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- Compressed Video Quality Enhancement: Classifying and Benchmarking over Standards
- BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals
- Local Information Matters: A Rethink of Crowd Counting
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
- mRadNet: A Compact Radar Object Detector with MetaFormer
- Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
- OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
- A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Live(r) Die: Predicting Survival in Colorectal Liver Metastasis
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
- First-order State Space Model for Lightweight Image Super-resolution
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- RepViT-CXR: A Channel Replication Strategy for Vision Transformers in Chest X-ray Tuberculosis and Pneumonia Classification
- Apollo: A Posteriori Label-Only Membership Inference Attack Towards Machine Unlearning
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- SEEC: Segmentation-Assisted Multi-Entropy Models for Learned Lossless Image Compression
- SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
- DR-CircuitGNN: Training Acceleration of Heterogeneous Circuit Graph Neural Network on GPUs
- H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- Hybrid Swin Attention Networks for Simultaneously Low-Dose PET and CT Denoising
- Integrated Detection and Tracking Based on Radar Range-Doppler Feature
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- Learning spatially structured open quantum dynamics with regional-attention transformers
- MRD-LiNet: A Novel Lightweight Hybrid CNN with Gradient-Guided Unlearning for Improved Drought Stress Identification
- IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
- AIM 2025 Challenge on High FPS Motion Deblurring: Methods and Results
- Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
- DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
- Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
- A brain-inspired paradigm for scalable quantum vision
- SpecSwin3D: Generating Hyperspectral Imagery from Multispectral Data via Transformer Networks
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- HyPINO: Multi-Physics Neural Operators via HyperPINNs and the Method of Manufactured Solutions
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Toward Accessible Dermatology: Skin Lesion Classification Using Deep Learning Models on Mobile-Acquired Images
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
- Robust End-to-End FSO Transmission with Joint Coding Modulation and BiLSTM-Based Channel Modeling under Atmospheric Turbulence
- RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
- Unsupervised Instance Segmentation with Superpixels
- LGBP-OrgaNet: Learnable Gaussian Band Pass Fusion of CNN and Transformer Features for Robust Organoid Segmentation and Tracking
- Gradient Estimation Methods of Approximate Multipliers for High-Accuracy Retraining of Deep Learning Models
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- RiverScope: High-Resolution River Masking Dataset
- Vision encoders should be image size agnostic and task driven
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- SpectMamba: Integrating Frequency and State Space Models for Enhanced Medical Image Detection
- Predicting prognosis of light-chain cardiac amyloidosis by magnetic resonance imaging and deep learning
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- DCA: Graph-Guided Deep Embedding Clustering for Brain Atlases
- Exploring Over-stationarization in Deep Learning-based Bus/Tram Arrival Time Prediction: Analysis and Non-stationary Effect Recovery
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- Satellite Image Utilization for Dehazing with Swin Transformer-Hybrid U-Net and Watershed loss
- Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition
- Encoder-Only Image Registration
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- FW-GAN: Frequency-Driven Handwriting Synthesis with Wave-Modulated MLP Generator
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- Prediction of Distant Metastasis in Head and Neck Cancer Patients Using Tumor and Peritumoral Multi-Modal Deep Learning
- Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection
- Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
- CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
- STGAtt: A Spatial-Temporal Unified Graph Attention Network for Traffic Flow Forecasting
- Foundation Models for Cross-Domain EEG Analysis Application: A Survey
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
- Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- Gradient Rectification for Robust Calibration under Distribution Shift
- From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Machine-learning competition to grade EEG background patterns in newborns with hypoxic-ischaemic encephalopathy
- Autoregressive Universal Video Segmentation Model
- Random forest-based out-of-distribution detection for robust lung cancer segmentation
- Can we make NeRF-based visual localization privacy-preserving?
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- Huracan: A skillful end-to-end data-driven system for ensemble data assimilation and weather prediction
- VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Signals vs. Videos: Advancing Motion Intention Recognition for Human-Robot Collaboration in Construction
- D3FNet: A Differential Attention Fusion Network for Fine-Grained Road Structure Extraction in Remote Perception Systems
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
- The Loupe: A Plug-and-Play Attention Module for Amplifying Discriminative Features in Vision Transformers
- You Only Pose Once: A Minimalist's Detection Transformer for Monocular RGB Category-level 9D Multi-Object Pose Estimation
- Reliable Smoke Detection via Optical Flow-Guided Feature Fusion and Transformer-Based Uncertainty Modeling
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
- MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
- TCFNet: Bidirectional face-bone transformation via a Transformer-based coarse-to-fine point movement network
- GasTwinFormer: A Hybrid Vision Transformer for Livestock Methane Emission Segmentation and Dietary Classification in Optical Gas Imaging
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- Communication-Efficient Federated Learning with Adaptive Number of Participants
- A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
- AIM 2025 challenge on Inverse Tone Mapping Report: Methods and Results
- Wavy Transformer
- Generalization vs. Memorization in Autoregressive Deep Learning: Or, Examining Temporal Decay of Gradient Coherence
- FractMorph: A Fractional Fourier-Based Multi-Domain Transformer for Deformable Image Registration
- Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems
- Geometry-Aware Video Inpainting for Joint Headset Occlusion Removal and Face Reconstruction in Social XR
- HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
- Illusions in Humans and AI: How Visual Perception Aligns and Diverges
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- MedFormer: a data-driven model for forecasting the Mediterranean Sea
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- ENA: Efficient N-dimensional Attention
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- AIM: Amending Inherent Interpretability via Self-Supervised Masking
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Importance-Aware Robust Semantic Transmission for LEO Satellite-Ground Communication
- Subcortical Masks Generation in CT Images via Ensemble-Based Cross-Domain Label Transfer
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
- Fourier-Guided Attention Upsampling for Image Super-Resolution
- DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
- PSScreen: Partially Supervised Multiple Retinal Disease Screening
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Improving Learning of New Diseases through Knowledge-Enhanced Initialization for Federated Adapter Tuning
- Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Teaching LLMs to Speak Spectroscopy
- Hierarchical Graph Attention Network for No-Reference Omnidirectional Image Quality Assessment
- MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention
- NEURAL: Attention-Guided Pruning for Unified Multimodal Resource-Constrained Clinical Evaluation
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Decentralized Rank Scheduling for Energy-Constrained Multi-Task Federated Fine-Tuning in Edge-Assisted IoV Networks
- Learning Spatial Decay for Vision Transformers
- RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- AI-Driven Detection and Analysis of Handwriting on Seized Ivory: A Tool to Uncover Criminal Networks in the Illicit Wildlife Trade
- FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
- UltraLight Med-Vision Mamba for Classification of Neoplastic Progression in Tubular Adenomas
- Automated Charge Transition Detection in Quantum Dot Charge Stability Diagrams
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- PADReg: Physics-Aware Deformable Registration Guided by Contact Force for Ultrasound Sequences
- QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Scaling Learned Image Compression Models up to 1 Billion
- Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
- SelfHVD: Self-Supervised Handheld Video Deblurring
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
- CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- Vision Generalist Model: A Survey
- MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- ∇NABLA: Neighborhood Adaptive Block-Level Attention
- SparseC-AFM: a deep learning method for fast and accurate characterization of MoS2 with C-AFM
- Text-guided Visual Prompt DINO for Generic Segmentation
- UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
- Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
- Hybrid(Transformer+CNN)-based Polyp Segmentation
- Lightweight Quad Bayer HybridEVS Demosaicing via State Space Augmented Cross-Attention
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- SMOL-MapSeg: Show Me One Label as prompt
- Deformable Attention Graph Representation Learning for Histopathology Whole Slide Image Analysis
- CT-GRAPH: Hierarchical Graph Attention Network for Anatomy-Guided CT Report Generation
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
- A Neural Conditional Random Field Model Using Deep Features and Learnable Functions for End-to-End MRI Prostate Zonal Segmentation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Unified modality separation: A vision-language framework for unsupervised domain adaptation
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image Compression
- Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
- Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- Benchmarking pig detection and tracking under diverse and challenging conditions
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
- Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- Benchmarking Foundation Models for Mitotic Figure Classification
- Efficient Inter-Task Attention for Multitask Transformer Models
- VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- A2Mamba: Attention-augmented State Space Models for Visual Recognition
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Deeper Inside Deep ViT
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
- Towards Globally Predictable k-Space Interpolation: A White-box Transformer Approach
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
- Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
- AttZoom: Attention Zoom for Better Visual Features
- Hidden Dynamics of Massive Activations in Transformer Training
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation
- Monocular Depth Estimation with Global-Aware Discretization and Local Context Modeling
- SSFMamba: Symmetry-driven Spatial-Frequency Feature Fusion for 3D Medical Image Segmentation
- CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
- Rethinking Transparent Object Grasping: Depth Completion with Monocular Depth Estimation and Instance Mask
- TRUDI and TITUS: A Multi-Perspective Dataset and A Three-Stage Recognition System for Transportation Unit Identification
- DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal
- S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
- Context Guided Transformer Entropy Modeling for Video Compression
- Large Kernel MedNeXt for Breast Tumor Segmentation and Self-Normalizing Network for pCR Classification in Magnetic Resonance Images
- EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
- DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- CGCCE-Net:Change-Guided Cross Correlation Enhancement Network for Remote Sensing Building Change Detection
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Self-Navigated Residual Mamba for Universal Industrial Anomaly Detection
- LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Skip priors and add graph-based anatomical information, for point-based Couinaud segmentation
- A Full-Stage Refined Proposal Algorithm for Suppressing False Positives in Two-Stage CNN-Based Detection Methods
- Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network
- Synthetic Data Matters: Re-training with Geo-typical Synthetic Labels for Building Detection
- SWAN: Synergistic Wavelet-Attention Network for Infrared Small Target Detection
- RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification
- Conquering High Packet-Loss Erasure: MoE Swin Transformer-Based Video Semantic Communication
- Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- A Framework Combining 3D CNN and Transformer for Video-Based Behavior Recognition
- Flow Matching for Probabilistic Learning of Dynamical Systems from Missing or Noisy Data
- Disrupting Semantic and Abstract Features for Better Adversarial Transferability
- OmniUnet: A Multimodal Network for Unstructured Terrain Segmentation on Planetary Rovers Using RGB, Depth, and Thermal Imagery
- DBLP: Noise Bridge Consistency Distillation For Efficient And Reliable Adversarial Purification
- Representation Shift: Unifying Token Compression with FlashAttention
- VQ-DeepISC: Vector Quantized-Enabled Digital Semantic Communication with Channel Adaptive Image Transmission
- Multimodal Referring Segmentation: A Survey
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
- MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space Model
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Adjustable Spatio-Spectral Hyperspectral Image Compression Network
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- Efficient Face Image Quality Assessment via Self-training and Knowledge Distillation
- Visual-Language Model Knowledge Distillation Method for Image Quality Assessment
- Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors
- DACA-Net: A Degradation-Aware Conditional Diffusion Network for Underwater Image Enhancement
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
- Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- Whole-brain Transferable Representations from Large-Scale fMRI Data Improve Task-Evoked Brain Activity Decoding
- MSQ: Memory-Efficient Bit Sparsification Quantization
- Learning from Heterogeneous Structural MRI via Collaborative Domain Adaptation for Late-Life Depression Assessment
- Vision-Language Cross-Attention for Real-Time Autonomous Driving
- From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
- Spatial-Temporal-Spectral Mamba with Sparse Deformable Token Sequence for Enhanced MODIS Time Series Classification
- Brain Tumor Segmentation in Sub-Sahara Africa with Advanced Transformer and ConvNet Methods: Fine-Tuning, Data Mixing and Ensembling
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Color as the Impetus: Transforming Few-Shot Learner
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning
- Staining and locking computer vision models without retraining
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
- SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
- Cross-Architecture Distillation Made Simple with Redundancy Suppression
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Towards White-Box Deep Wireless Sensing
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
- Regularizing Subspace Redundancy of Low-Rank Adaptation
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
- ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- Few-Shot Object Detection via Spatial-Channel State Space Model
- EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control
- AnimalClue: Recognizing Animals by their Traces
- Local2Global query Alignment for Video Instance Segmentation
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Extreme Solar Storm Reveals Causal Interactions in Space Weather
- Dual-Stream Global-Local Feature Collaborative Representation Network for Scene Classification of Mining Area
- SkinDualGen: Prompt-Driven Diffusion for Simultaneous Image-Mask Generation in Skin Lesions
- Cross-Domain Few-Shot Learning with Coalescent Projections and Latent Space Reservation
- Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
- Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
- Taming Domain Shift in Multi-source CT-Scan Classification via Input-Space Standardization
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- Large Language Model Agent for Structural Drawing Generation Using ReAct Prompt Engineering and Retrieval Augmented Generation
- Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
- MambaVesselNet++: A Hybrid CNN-Mamba Architecture for Medical Image Segmentation
- Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and beyond
- SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
- Modality Agnostic Efficient Long Range Encoder
- WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design
- MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
- PatchTraj: Unified Time-Frequency Representation Learning via Dynamic Patches for Trajectory Prediction
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Self-training Framework for Semi-supervised Pulmonary Vessel Segmentation and Its Application in COPD
- MedIQA: A Scalable Foundation Model for Prompt-Driven Medical Image Quality Assessment
- Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments
- Deformable Convolution Module with Globally Learned Relative Offsets for Fundus Vessel Segmentation
- VB-Mitigator: An Open-source Framework for Evaluating and Advancing Visual Bias Mitigation
- Exploiting Gaussian Agnostic Representation Learning with Diffusion Priors for Enhanced Infrared Small Target Detection
- Differential-UMamba: Rethinking Tumor Segmentation Under Limited Data Scenarios
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- Explaining How Visual, Textual and Multimodal Encoders Share Concepts
- Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
- Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss
- Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)
- DiNAT-IR: Exploring Dilated Neighborhood Attention for High-Quality Image Restoration
- Open-set Cross Modal Generalization via Multimodal Unified Representation
- Region-aware Depth Scale Adaptation with Sparse Measurements
- Training Self-Supervised Depth Completion Using Sparse Measurements and a Single Image
- Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
- BloomSight: An ultra-high-frequency phenotyping framework for diurnal flowering dynamics in japonica and indica rice to enable genetic dissection and hybrid-breeding applications
- edge-SR: Super-Resolution For The Masses
- DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
- Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
- Open-World Entity Segmentation
- UGPL: Uncertainty-Guided Progressive Learning for Evidence-Based Classification in Computed Tomography
- Translational application of a self-organized deep feature engineering pipeline for non-invasive pulmonary hypertension classification from routine chest radiographs
- A pulmonary nodule is worth 8 × 8 × 8 words: Computed tomography-based 3D vision transformer predicts early-stage high-grade lung adenocarcinoma of micropapillary and/or solid subtypes
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation
- Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
- On the Interaction of Compressibility and Adversarial Robustness
- Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based Priors
- Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
- Swin-TUNA : A Novel PEFT Approach for Accurate Food Image Segmentation
- Efficient Burst Super-Resolution with One-step Diffusion
- Global Modeling Matters: A Fast, Lightweight and Effective Baseline for Efficient Image Restoration
- ParallelTime: Dynamically Weighting the Balance of Short- and Long-Term Temporal Dependencies
- DOOMGAN:High-Fidelity Dynamic Identity Obfuscation Ocular Generative Morphing
- SRMambaV2: Biomimetic Attention for Sparse Point Cloud Upsampling in Autonomous Driving
- Leveraging Pathology Foundation Models for Panoptic Segmentation of Melanoma in H&E Images
- Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
- Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification
- Few-Shot Learning in Video and 3D Object Detection: A Survey
- Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- EPSilon: Efficient Point Sampling for Lightening of Hybrid-based 3D Avatar Generation
- Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
- HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery
- Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
- ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
- DenseSR: Image Shadow Removal as Dense Prediction
- Variance-Based Pruning for Accelerating and Compressing Trained Networks
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- AtrousMamaba: An Atrous-Window Scanning Visual State Space Model for Remote Sensing Change Detection
- AMMNet: An Asymmetric Multi-Modal Network for Remote Sensing Semantic Segmentation
- A Hybrid CNN-VSSM model for Multi-View, Multi-Task Mammography Analysis: Robust Diagnosis with Attention-Based Fusion
- Deep Image Reconstruction for Background Subtraction in Heavy-Ion Collisions
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- Compact Vision Transformer by Reduction of Kernel Complexity
- DASViT: Differentiable Architecture Search for Vision Transformer
- Learning-Based Interface for Semantic Communication with Bit Importance Awareness
- Change of Thought: Adaptive Test-Time Computation
- A Deep-Learning Framework for Land-Sliding Classification from Remote Sensing Image
- SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection
- MVA 2025 Small Multi-Object Tracking for Spotting Birds Challenge: Dataset, Methods, and Results
- Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
- Canonical Latent Representations in Conditional Diffusion Models
- Dataset Ownership Verification for Pre-trained Masked Models
- Frequency-Dynamic Attention Modulation for Dense Prediction
- SAMST: A Transformer framework based on SAM pseudo label filtering for remote sensing semi-supervised semantic segmentation
- A Survey of Deep Learning for Geometry Problem Solving
- CompressedVQA-HDR: Generalized Full-reference and No-reference Quality Assessment Models for Compressed High Dynamic Range Videos
- GLOMIA-Pro: A Generalizable Longitudinal Medical Image Analysis Framework for Disease Progression Prediction
- Spatial Frequency Modulation for Semantic Segmentation
- Image-Based Multi-Survey Classification of Light Curves with a Pre-Trained Vision Transformer
- MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
- Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
- CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- MAC: An Efficient Gradient Preconditioning using Mean Activation Approximated Curvature
- A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging
- Hyperspectral Image Classification via Transformer-based Spectral-Spatial Attention Decoupling and Adaptive Gating
- Bias Analysis in Unconditional Image Generative Models
- Plug-and-play linear attention with provable guarantees for training-free image restoration
- SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics
- CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
- Multimodal Representation Alignment for Cross-modal Information Retrieval
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
- Graph Aggregation Prototype Learning for Semantic Change Detection in Remote Sensing
- InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba
- Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- CodeBrain: Bridging Decoupled Tokenizer and Multi-Scale Architecture for EEG Foundation Model
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
- DepViT-CAD: Deployable Vision Transformer-Based Cancer Diagnosis in Histopathology
- Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
- Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
- Leveraging Swin Transformer for enhanced diagnosis of Alzheimer's disease using multi-shell diffusion MRI
- Foundation Models in Medical Imaging: A Review and Outlook
- Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
- A Transfer Learning-Based Method for Water Body Segmentation in Remote Sensing Imagery: A Case Study of the Zhada Tulin Area
- KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection
- MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
- MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding
- QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
- PPJudge: Towards Human-Aligned Assessment of Artistic Painting Process
- A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
- Disentanglement and Assessment of Shortcuts in Ophthalmological Retinal Imaging Exams
- Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
- Controllable Patching for Compute-Adaptive Surrogate Modeling of Partial Differential Equations
- Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning
- RoHOI: Robustness Benchmark for Human-Object Interaction Detection
- Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
- BioAnalyst: A Foundation Model for Biodiversity
- Model Parallelism With Subnetwork Data Parallelism
- HieraRS: A Hierarchical Segmentation Paradigm for Remote Sensing Enabling Multi-Granularity Interpretation and Cross-Domain Transfer
- L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
- Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
- Critical dynamics governs deep learning
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
- An Adaptive Volatility-based Learning Rate Scheduler
- PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
- HNOSeg-XS: Extremely Small Hartley Neural Operator for Efficient and Resolution-Robust 3D Image Segmentation
- Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
- Synergistic Prompting for Robust Visual Recognition with Missing Modalities
- Bridging the gap in FER: addressing age bias in deep learning
- Advancing Medical Image Segmentation via Self-supervised Instance-adaptive Prototype Learning
- Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
- Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning
- PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and Consistency
- SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples
- Visual Instance-aware Prompt Tuning
- Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning
- IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation Optimization
- StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
- PointVDP: Learning View-Dependent Projection by Fireworks Rays for 3D Point Cloud Segmentation
- Omni-Fusion of Spatial and Spectral for Hyperspectral Image Segmentation
- Capturing Stable HDR Videos Using a Dual-Camera System
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
- Mammo-Clustering: Context Clustering based Multi-view Tri Level Information Fusion for Lesion Location and Classification in Mammography
- ScoreAdv: Score-based Targeted Generation of Natural Adversarial Examples via Diffusion Models
- Fair Domain Generalization: An Information-Theoretic View
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning
- Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset
- RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
- DREAM: Document Reconstruction via End-to-end Autoregressive Model
- Ampere: Communication-Efficient and High-Accuracy Split Federated Learning
- Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
- Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation
- Towards fair decentralized benchmarking of healthcare AI algorithms with the Federated Tumor Segmentation (FeTS) challenge
- VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
- A Federated Learning-based Lightweight Network with Zero Trust for UAV Authentication
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- DARIL: When Imitation Learning outperforms Reinforcement Learning in Surgical Action Planning
- MVNet: Hyperspectral Remote Sensing Image Classification Based on Hybrid Mamba-Transformer Vision Backbone Architecture
- Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration
- Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
- ViTaL: A Multimodality Dataset and Benchmark for Multi-pathological Ovarian Tumor Recognition
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- Early Convolutions Help Transformers See Better
- Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
- Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs
- Zero Memory Overhead Approach for Protecting Vision Transformer Parameters
- Be the Change You Want to See: Revisiting Remote Sensing Change Detection Practices
- CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection
- NOVO: Unlearning-Compliant Vision Transformers
- Dual-frequency Selected Knowledge Distillation with Statistical-based Sample Rectification for PolSAR Image Classification
- StreamDiT: Real-Time Streaming Text-to-Video Generation
- Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
- Temporal Window Smoothing of Exogenous Variables for Improved Time Series Prediction
- HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
- Time-Masked Transformers with Lightweight Test-Time Adaptation for Neural Speech Decoding
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
- Transformer-based EEG Decoding: A Survey
- High-Fidelity Differential-information Driven Binary Vision Transformer
- LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection
- Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
- STEM Diffraction Pattern Analysis with Deep Learning Networks
- Modulate and Reconstruct: Learning Hyperspectral Imaging from Misaligned Smartphone Views
- SSL4SAR: Self-Supervised Learning for Glacier Calving Front Extraction from SAR Imagery
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- DaiFu: In-Situ Crash Recovery for Deep Learning Systems
- Enhancing Multi-Exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning
- Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
- EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
- POST: Photonic Swin Transformer for Automated and Efficient Prediction of PCSEL
- Learning an Ensemble Token from Task-driven Priors in Facial Analysis
- PhyUnfold-Net: Advancing Remote Sensing Change Detection with Physics-Guided Deep Unfolding
- evMLP: An Efficient Event-Driven MLP Architecture for Vision
- Multi Source COVID-19 Detection via Kernel-Density-based Slice Sampling
- Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
- Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
- Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
- Rectifying Magnitude Neglect in Linear Attention
- World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
- Few-shot Classification as Multi-instance Verification: Effective Backbone-agnostic Transfer across Domains
- MedDiff-FT: Data-Efficient Diffusion Model Fine-tuning with Structural Guidance for Controllable Medical Image Synthesis
- Customizable ROI-Based Deep Image Compression
- Out-of-Distribution Detection with Adaptive Top-K Logits Integration
- CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation
- Do Protein Transformers Have Biological Intelligence?
- Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
- MAMBO: High-Resolution Generative Approach for Mammography Images
- Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections
- SelvaBox: A high-resolution dataset for tropical tree crown detection
- FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
- Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios
- Visual Textualization for Image Prompted Object Detection
- Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
- MedSAM-CA: A CNN-Augmented ViT with Attention-Enhanced Multi-Scale Fusion for Medical Image Segmentation
- Pruning by Block Benefit: Exploring the Properties of Vision Transformer Blocks during Domain Adaptation
- Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training Acceleration
- OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- Sample Margin-Aware Recalibration of Temperature Scaling
- Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
- Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
- From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
- Three-dimensional end-to-end deep learning for brain MRI analysis
- Low-latency vision transformers via large-scale multi-head attention
- FD-DiT: Frequency Domain-Directed Diffusion Transformer for Low-Dose CT Reconstruction
- PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
- A Hierarchical Slice Attention Network for Appendicitis Classification in 3D CT Scans
- Detecting What Matters: A Novel Approach for Out-of-Distribution 3D Object Detection in Autonomous Vehicles
- DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
- Ensemble-Based Survival Models with the Self-Attended Beran Estimator Predictions
- Attention to the Burstiness in Visual Prompt Tuning!
- ICME 2025 Generalizable HDR and SDR Video Quality Measurement Grand Challenge
- SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
- YM-WML: A new Yolo-based segmentation Model with Weighted Multi-class Loss for medical imaging
- Are Fast Methods Stable in Adversarially Robust Transfer Learning?
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- Improving Token-based Object Detection with Video
- Design and Evaluation of Deep Learning-Based Dual-Spectrum Image Fusion Methods
- Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing Images
- End-to-End RGB-IR Joint Image Compression With Channel-wise Cross-modality Entropy Model
- Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
- Benchmarking Deep Learning and Vision Foundation Models for Atypical vs. Normal Mitosis Classification with Cross-Dataset Evaluation
- Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
- Pushing Trade-Off Boundaries: Compact yet Effective Remote Sensing Change Detection
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
- RL-Selector: Reinforcement Learning-Guided Data Selection via Redundancy Assessment
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
- Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning
- VisionGuard: Synergistic Framework for Helmet Violation Detection
- M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery Localization
- FastRef:Fast Prototype Refinement for Few-Shot Industrial Anomaly Detection
- Norm×Direction: Restoring the Missing Query Norm in Vision Linear Attention
- InvZW: Invariant Feature Learning via Noise-Adversarial Training for Robust Image Zero-Watermarking
- EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis
- Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-Supervision, Mixture of Experts and Multi-Task Integration
- ViFusionTST: Deep Fusion of Time-Series Image Representations from Load Signals for Early Bed-Exit Prediction
- U-R-VEDA: Integrating UNET, Residual Links, Edge and Dual Attention, and Vision Transformer for Accurate Semantic Segmentation of CMRs
- MS-IQA: A Multi-Scale Feature Fusion Network for PET/CT Image Quality Assessment
- A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
- Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- HMSViT: A Hierarchical Masked Self-Supervised Vision Transformer for Corneal Nerve Segmentation and Diabetic Neuropathy Diagnosis
- A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation
- Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents
- Automated Image Recognition Framework
- SimpleGVR: A Simple Baseline for Latent-Cascaded Video Super-Resolution
- Reconsidering Explicit Longitudinal Mammography Alignment for Enhanced Breast Cancer Risk Prediction
- OpenWildlife: Open-Vocabulary Multi-Species Wildlife Detector for Geographically-Diverse Aerial Imagery
- LKA: Large Kernel Adapter for Enhanced Medical Image Classification
- Focus Your Attention: Towards Data-Intuitive Lightweight Vision Transformers
- Including Semantic Information via Word Embeddings for Skeleton-based Action Recognition
- DPFormer: Dynamic Prompt Transformer for Continual Learning
- Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
- FAMSeg: Fetal Femur and Cranial Ultrasound Segmentation Using Feature-Aware Attention and Mamba Enhancement
- Transforming H&E images into IHC: A Variance-Penalized GAN for Precision Oncology
- Open Set Recognition for Endoscopic Image Classification: A Deep Learning Approach on the Kvasir Dataset
- DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
- Improving Black-Box Generative Attacks via Generator Semantic Consistency
- Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
- Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
- From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
- OSDMamba: Enhancing Oil Spill Detection from Remote Sensing Images Using Selective State Space Model
- h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective
- Cloud-Aware SAR Fusion for Enhanced Optical Sensing in Space Missions
- NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
- SynDaCaTE: A Synthetic Dataset For Evaluating Part-Whole Hierarchical Inference
- MTSIC: Multi-stage Transformer-based GAN for Spectral Infrared Image Colorization
- Trans2-CBCT: A Dual-Transformer Framework for Sparse-View CBCT Reconstruction
- From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge
- Multi-label Scene Classification for Autonomous Vehicles: Acquiring and Accumulating Knowledge from Diverse Datasets
- Universal Music Representations? Evaluating Foundation Models on World Music Corpora
- Relaxed syntax modeling in Transformers for future-proof license plate recognition
- Unsupervised Image Super-Resolution Reconstruction Based on Real-World Degradation Patterns
- Noise-Informed Diffusion-Generated Image Detection with Anomaly Attention
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- Single-step Diffusion for Image Compression at Ultra-Low Bitrates
- Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
- AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
- Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments
- Polyline Path Masked Attention for Vision Transformer
- Learning Multi-scale Spatial-frequency Features for Image Denoising
- Reliable Few-shot Learning under Dual Noises
- Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
- Echo-DND: A dual noise diffusion model for robust and precise left ventricle segmentation in echocardiography
- NTIRE 2025 Image Shadow Removal Challenge Report
- Mondrian: Transformer Operators via Domain Decomposition
- OpenPath: Open-Set Active Learning for Pathology Image Classification via Pre-trained Vision-Language Models
- Retrospective Memory for Camouflaged Object Detection
- Partitioning for Intrinsic Model Inversion Resistance in Collaborative Inference
- MapFM: Foundation Model-Driven HD Mapping with Multi-Task Contextual Learning
- Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning
- Earth Observation Foundation Model PhilEO: Pretraining on the MajorTOM and FastTOM Datasets
- ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views
- Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems
- SCISSOR: Mitigating Semantic Bias through Cluster-Aware Siamese Networks for Robust Classification
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
- Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
- SeqPE: Transformer with Sequential Position Encoding
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- Evolution of ReID: From Early Methods to LLM Integration
- COME: Adding Scene-Centric Forecasting Control to Occupancy World Model
- Hierarchical Multi-Positive Contrastive Learning for Patent Image Retrieval
- Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
- FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception
- Overcoming Occlusions in the Wild: A Multi-Task Age Head Approach to Age Estimation
- ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs
- Learning Event Completeness for Weakly Supervised Video Anomaly Detection
- Distributional Training Data Attribution: What do Influence Functions Sample?
- Boosting Adversarial Transferability via Commonality-Oriented Gradient Optimization
- Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
- Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models
- Unleashing Diffusion and State Space Models for Medical Image Segmentation
- Combining Self-attention and Dilation Convolutional for Semantic Segmentation of Coal Maceral Groups
- Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation
- DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
- Efficient Star Distillation Attention Network for Lightweight Image Super-Resolution
- FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
- GPLQ: A General, Practical, and Lightning QAT Method for Vision Transformers
- Voxel-Level Brain States Prediction Using Swin Transformer
- ReStNet: A Reusable & Stitchable Network for Dynamic Adaptation on IoT Devices
- Polar Hierarchical Mamba: Towards Streaming LiDAR Object Detection with Point Clouds as Egocentric Sequences
- Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
- Video Frame Interpolation Transformer
- PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis
- FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
- DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
- J-DDL: Surface Damage Detection and Localization System for Fighter Aircraft
- SLICK: Selective Localization and Instance Calibration for Knowledge-Enhanced Car Damage Segmentation in Automotive Insurance
- ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation
- Don't Pay Attention
- California Crop Yield Benchmark: Combining Satellite Image, Climate, Evapotranspiration, and Soil Data Layers for County-Level Yield Forecasting of Over 70 Crops
- Detecção da Psoríase Utilizando Visão Computacional: Uma Abordagem Comparativa Entre CNNs e Vision Transformers
- The Four Color Theorem for Cell Instance Segmentation
- Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks
- SatelliteFormula: Multi-Modal Symbolic Regression from Remote Sensing Imagery for Physics Discovery
- LaDEEP: A Deep Learning-based Surrogate Model for Large Deformation of Elastic-Plastic Solids
- Adaptive Contextual Embedding for Robust Far-View Borehole Detection
- Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
- DermaCon-IN: A Multi-concept Annotated Dermatological Image Dataset of Indian Skin Disorders for Clinical AI Research
- Tensor-to-Tensor Models with Fast Iterated Sum Features
- Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
- MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation
- Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
- GS4: Generalizable Sparse Splatting Semantic SLAM
- WoundAIssist: A Patient-Centered Mobile App for AI-Assisted Wound Care With Physicians in the Loop
- SSH-Net: A Self-Supervised and Hybrid Network for Noisy Image Watermark Removal
- Sample-Specific Noise Injection For Diffusion-Based Adversarial Purification
- Self-supervised One-Stage Learning for RF-based Multi-Person Pose Estimation
- F2T2-HiT: A U-Shaped FFT Transformer and Hierarchical Transformer for Reflection Removal
- DualX-VSR: Dual Axial Spatial×Temporal Transformer for Real-World Video Super-Resolution without Motion Compensation
- APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
- HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
- A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
- KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance Products
- Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
- Towards Comprehensive Monocular Depth Estimation: Multiple Heads Are Better Than One
- SwinLSTM Autoencoder for Temporal-Spatial-Frequency Domain CSI Compression in Massive MIMO Systems
- TS-SNN: Temporal Shift Module for Spiking Neural Networks
- Vision Graph Prompting via Semantic Low-Rank Decomposition
- Image Restoration via Multi-domain Learning
- Lightweight RGB-D Salient Object Detection from a Speed-Accuracy Tradeoff Perspective
- Tetrahedron-Net for Medical Image Registration
- WDMamba: When Wavelet Degradation Prior Meets Vision Mamba for Image Dehazing
- ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
- Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
- ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
- Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?
- CM1 -- A Dataset for Evaluating Few-Shot Information Extraction with Large Vision Language Models
- False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
- Enhanced SCanNet with CBAM and Dice Loss for Semantic Change Detection
- AI Agent Behavioral Science
- BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
- SAAT: Synergistic Alternating Aggregation Transformer for Image Super-Resolution
- Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector
- ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
- A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark
- Res-MoCoDiff: Residual-guided diffusion models for motion artifact correction in brain MRI
- FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation
- FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition
- Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?
- ControlMambaIR: Conditional Controls with State-Space Model for Image Restoration
- BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations
- RoadFormer : Local-Global Feature Fusion for Road Surface Classification in Autonomous Driving
- Controllable Human-centric Keyframe Interpolation with Generative Prior
- ConMamba: Contrastive Vision Mamba for Plant Disease Detection
- OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction
- Towards In-the-wild 3D Plane Reconstruction from a Single Image
- AutoCMR: an automated pipeline for 3D CMR acquisition, reconstruction, and analysis.
- Random Registers for Cross-Domain Few-Shot Learning
- Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
- Self-Supervised Spatial Correspondence Across Modalities
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
- InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
- Image Recognition with Online Lightweight Vision Transformer: A Survey
- Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning
- No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
- Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
- AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
- Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges
- Quotient Network -- A Network Similar to ResNet but Learning Quotients
- GRAM: Spatial general-purpose audio representation models for real-world applications
- Low-Complexity Patch-Based No-Reference Point Cloud Quality Metric Exploiting Weighted Structure and Texture Features
- CAPAA: Classifier-Agnostic Projector-Based Adversarial Attack
- LoRA as a Flexible Framework for Securing Large Vision Systems
- iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection
- Test-time Vocabulary Adaptation for Language-driven Object Detection
- DCS-ST for Classification of Breast Cancer Histopathology Images with Limited Annotations
- Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
- PathGene: Benchmarking Driver Gene Mutations and Exon Prediction Using Multicenter Lung Cancer Histopathology Image Dataset
- Deformable Attention Mechanisms Applied to Object Detection, case of Remote Sensing
- ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation
- S3CE-Net: Spike-guided Spatiotemporal Semantic Coupling and Expansion Network for Long Sequence Event Re-Identification
- Beyond the LUMIR challenge: The pathway to foundational registration models
- Towards Unified Modeling in Federated Multi-Task Learning via Subspace Decoupling
- Deep learning-derived arterial input function for dynamic brain PET
- PDE-Transformer: Efficient and Versatile Transformers for Physics Simulations
- 50 Years of Automated Face Recognition
- DLiPath: A Benchmark for the Comprehensive Assessment of Donor Liver Based on Histopathological Image Dataset
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
- ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer
- PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
- MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
- A New Deep-learning-Based Approach For mRNA Optimization: High Fidelity, Computation Efficiency, and Multiple Optimization Factors
- HyperPointFormer: Multimodal Fusion in 3D Space with Dual-Branch Cross-Attention Transformers
- Generalizable Video Quality Assessment via Weak-to-Strong Learning
- DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment
- Database-Agnostic Gait Enrollment using SetTransformers
- Advances in Automated Fetal Brain MRI Segmentation and Biometry: Insights from the FeTA 2024 Challenge
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
- From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?
- Revisiting Reweighted Risk for Calibration: AURC, Focal, and Inverse Focal Loss
- What is the Right Embedding Space for Contrastive Learning in REC?
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape
- Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis
- The Resurrection of the ReLU
- AquaMonitor: A multimodal multi-view image sequence dataset for real-life aquatic invertebrate biodiversity monitoring
- Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
- RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination
- YH-MINER: Multimodal Intelligent System for Natural Ecological Reef Metric Extraction
- Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
- Towards Scalable Language-Image Pre-training for 3D Medical Imaging
- Synonymous Variational Inference for Perceptual Image Compression
- S2AFormer: Strip Self-Attention for Efficient Vision Transformer
- RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
- Taming Transformer Without Using Learning Rate Warmup
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- STA-Risk: A Deep Dive of Spatio-Temporal Asymmetries for Breast Cancer Risk Prediction
- AgriFM: A Multi-source Temporal Remote Sensing Foundation Model for Crop Mapping
- Occlusion Boundary and Depth: Mutual Enhancement via Multi-Task Learning
- Token Coordinated Prompt Attention is Needed for Visual Prompting
- NeuralOM: Neural Ocean Model for Subseasonal-to-Seasonal Simulation
- DSOcc: Leveraging Depth Awareness and Semantic Aid to Boost Camera-Based 3D Semantic Occupancy Prediction
- In Context Learning with Vision Transformers: Case Study
- SageAttention2++: A More Efficient Implementation of SageAttention2
- Object Concepts Emerge from Motion
- Open-Det: An Efficient Learning Framework for Open-Ended Detection
- Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
- Vision Transformers with Self-Distilled Registers
- MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation
- WDMIR: Wavelet-Driven Multimodal Intent Recognition
- BaryIR: Learning Multi-Source Unified Representation in Continuous Barycenter Space for Generalizable All-in-One Image Restoration
- QwT-v2: Practical, Effective and Efficient Post-Training Quantization
- HTMNet: A Hybrid Network with Transformer-Mamba Bottleneck Multimodal Fusion for Transparent and Reflective Objects Depth Completion
- VisAlgae 2023: A Dataset and Challenge for Algae Detection in Microscopy Images
- Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
- NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
- The Missing Point in Vision Transformers for Universal Image Segmentation
- HAODiff: Human-Aware One-Step Diffusion via Dual-Prompt Guidance
- CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge
- Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
- A Feature-level Bias Evaluation Framework for Facial Expression Recognition Models
- Adaptive Data-Resilient Multi-Modal Hierarchical Multi-Label Book Genre Identification
- Spurious Privacy Leakage in Neural Networks
- NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities
- Structured Initialization for Vision Transformers
- Locality-Aware Zero-Shot Human-Object Interaction Detection
- RGC-Bent: A Novel Dataset for Bent Radio Galaxy Classification
- OASIS: Optimized Lightweight Autoencoder System for Distributed In-Sensor computing
- Unaligned RGB Guided Hyperspectral Image Super-Resolution with Spatial-Spectral Concordance
- Smart Waste Management System for Makkah City using Artificial Intelligence and Internet of Things
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- Latent Mamba Operator for Partial Differential Equations
- A Smart Healthcare System for Monkeypox Skin Lesion Detection and Tracking
- MMET: A Multi-Input and Multi-Scale Transformer for Efficient PDEs Solving
- Reasoning Segmentation for Images and Videos: A Survey
- MLLMs are Deeply Affected by Modality Bias
- Asymmetric Duos: Sidekicks Improve Uncertainty
- Spiking Neural Networks Need High Frequency Information
- Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
- MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation
- Self-Organizing Visual Prototypes for Non-Parametric Representation Learning
- CarboFormer: A Lightweight Semantic Segmentation Architecture for Efficient Carbon Dioxide Detection Using Optical Gas Imaging
- Semantic Correspondence: Unified Benchmarking and a Strong Baseline
- GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Temporal Consistency Constrained Transferable Adversarial Attacks with Background Mixup for Action Recognition
- PEAR: Equal Area Weather Forecasting on the Sphere
- EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy
- PawPrint: Whose Footprints Are These? Identifying Animal Individuals by Their Footprints
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- From Flight to Insight: Semantic 3D Reconstruction for Aerial Inspection via Gaussian Splatting and Language-Guided Segmentation
- EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
- Repurposing Marigold for Zero-Shot Metric Depth Estimation via Defocus Blur Cues
- CENet: Context Enhancement Network for Medical Image Segmentation
- RemoteSAM: Towards Segment Anything for Earth Observation
- Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
- Training on Plausible Counterfactuals Removes Spurious Correlations
- AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
- NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
- TRAIL: Transferable Robust Adversarial Images via Latent diffusion
- Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention
- One-Step Diffusion-Based Image Compression with Semantic Distillation
- Native Segmentation Vision Transformers
- Improving Generalization in Heterogeneous Federated Continual Learning via Spatio-Temporal Gradient Matching with Prototypical Coreset
- SCOPE: Entanglement Frontier Escape for Source-Free Class Unlearning
- PCMamba: Physics-Informed Cross-Modal State Space Model for Dual-Camera Compressive Hyperspectral Imaging
- SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
- AuxDet: Auxiliary Metadata Matters for Omni-Domain Infrared Small Target Detection
- Time Tracker: Mixture-of-Experts-Enhanced Foundation Time Series Forecasting Model with Decoupled Training Pipelines
- Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
- Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
- Multi-View Projection for Unsupervised Domain Adaptation in 3D Semantic Segmentation
- Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning
- Loggia dei Lanzi: AI Thermography Enhancement Comparisons through 3D Photogrammetry
- Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNs
- SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
- Oral Imaging for Malocclusion Issues Assessments: OMNI Dataset, Deep Learning Baselines and Benchmarking
- GenFT: A Generative Parameter-Efficient Fine-Tuning Method for Pretrained Foundation Models
- Domain Adaptive Skin Lesion Classification via Conformal Ensemble of Vision Transformers
- UNet with Self-Adaptive Mamba-Like Attention and Causal-Resonance Learning for Medical Image Segmentation
- Know When to Abstain: Optimal Selective Classification with Likelihood Ratios
- Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
- Lossless Token Merging Even Without Fine-Tuning in Vision Transformers
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
- Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition
- A Unified Gradient-based Framework for Task-agnostic Continual Learning-Unlearning
- Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
- EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
- FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
- Multi-Channel Swin Transformer Framework for Bearing Remaining Useful Life Prediction
- SuperMapNet for Long-Range and High-Accuracy Vectorized HD Map Construction
- Towards Efficient Multi-Scale Deformable Attention on NPU
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- PiT: Progressive Diffusion Transformer
- Learning to Adapt to Position Bias in Vision Transformer Classifiers
- Attention-based clustering
- Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
- Aneumo: A Large-Scale Multimodal Aneurysm Dataset with Computational Fluid Dynamics Simulations and Deep Learning Benchmarks
- Rethinking Features-Fused-Pyramid-Neck for Object Detection
- Enhancing Transformers Through Conditioned Embedded Tokens
- Enhancing Channel-Independent Time Series Forecasting via Cross-Variate Patch Embedding
- Use as Many Surrogates as You Want: Selective Ensemble Attack to Unleash Transferability without Sacrificing Resource Efficiency
- Single Image Reflection Separation via Dual Prior Interaction Transformer
- AGI-Elo: How Far Are We From Mastering A Task?
- Mamba-Adaptor: State Space Model Adaptor for Visual Recognition
- MatPredict: a dataset and benchmark for learning material properties of diverse indoor objects
- Industrial Synthetic Segment Pre-training
- Dynamic Graph Induced Contour-aware Heat Conduction Network for Event-based Object Detection
- DD-Ranking: Rethinking the Evaluation of Dataset Distillation
- MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
- Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
- DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation
- CTLformer: A Hybrid Denoising Model Combining Convolutional Layers and Self-Attention for Enhanced CT Image Reconstruction
- Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object Detection
- TDFormer: A Top-Down Attention-Controlled Spiking Transformer
- MedVKAN: Efficient Feature Extraction with Mamba and KAN for Medical Image Segmentation
- SepPrune: Structured Pruning for Efficient Deep Speech Separation
- UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- FlashBias: Fast Computation of Attention with Bias
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
- MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
- ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
- Mask-Based Priors Are More Persistent than Query-Key Initializations
- Generalizable cardiac substructures segmentation from contrast and non-contrast CTs using pretrained transformers
- Assessing the Performance of Analog Training for Transfer Learning
- Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
- MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models
- WeGA: Weakly-Supervised Global-Local Affinity Learning Framework for Lymph Node Metastasis Prediction in Rectal Cancer
- From Preimage Search To Source-Grounded Feature Inversion
- IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation
- StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
- Hierarchical Surgical Robot Transformer (SRT-H): Imitation Learning for Autonomous Surgery
- Generating Full-field Evolution of Physical Dynamics from Irregular Sparse Observations
- A Multimodal Multi-Agent Framework for Radiology Report Generation
- FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration
- FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization
- AdaFortiTran: An Adaptive Transformer Model for Robust OFDM Channel Estimation
- A 2D Semantic-Aware Position Encoding for Vision Transformers
- TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series
- Enhancing Thyroid Cytology Diagnosis with RAG-Optimized LLMs and Pa-thology Foundation Models
- Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
- DHECA-SuperGaze: Dual Head-Eye Cross-Attention and Super-Resolution for Unconstrained Gaze Estimation
- FAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning
- CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
- AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
- DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
- Gameplay Highlights Generation
- A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
- DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection
- Autonomous Robotic Pruning in Orchards and Vineyards: a Review
- H3DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
- You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling with Gradient Shortcuts
- Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
- A Vision-Language Foundation Model for Leaf Disease Identification
- Technical Report for ICRA 2025 GOOSE 2D Semantic Segmentation Challenge: Leveraging Color Shift Correction, RoPE-Swin Backbone, and Quantile-based Label Denoising Strategy for Robust Outdoor Scene Understanding
- VALISENS: A Validated Innovative Multi-Sensor System for Cooperative Automated Driving
- Uni-AIMS: AI-Powered Microscopy Image Analysis
- Efficient and Robust Multidimensional Attention in Remote Physiological Sensing through Target Signal Constrained Factorization
- Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
- A Unified Pruning Framework for Vision Transformers
- Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
- Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
- A Lightweight Graph Transformer Network for Human Mesh Reconstruction from 2D Human Pose
- The Application of Deep Learning for Lymph Node Segmentation: A Systematic Review
- From AI Weather Prediction to Infrastructure Resilience: A Real-Time Correction-Downscaling Framework for Tropical Cyclone Impact Forecasting
- Predicting Diabetic Macular Edema Treatment Responses Using OCT: Dataset and Methods of APTOS Competition
- A review of advancements in low-light image enhancement using deep learning
- ProTCT: Projection quantification and fidelity constraint integrated deep reconstruction for Tangential CT
- Accurate and Efficient Multivariate Time Series Forecasting via Offline Clustering
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- DFEN: Dual Feature Equalization Network for Medical Image Segmentation
- The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction
- FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
- GaMNet: A Hybrid Network with Gabor Fusion and NMamba for Efficient 3D Glioma Segmentation
- Physics-Assisted and Topology-Informed Deep Learning for Weather Prediction
- A Simple Detector with Frame Dynamics is a Strong Tracker
- Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
- Augmented Deep Contexts for Spatially Embedded Video Coding
- How to build the best medical image segmentation algorithm using foundation models: a comprehensive empirical study with Segment Anything Model
- Cross-Branch Orthogonality for Improved Generalization in Face Deepfake Detection
- XtraLight-MedMamba for Classification of Neoplastic Tubular Adenomas
- Always Skip Attention
- Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
- Adversarial Robustness of Deep Learning Models for Inland Water Body Segmentation from SAR Images
- Scalable Machines with Intrinsic Higher Mental-State Dynamics
- PainFormer: a Vision Foundation Model for Automatic Pain Assessment
- T-Graph: Enhancing Sparse-view Camera Pose Estimation by Pairwise Translation Graph
- SemSpaceFL: A Collaborative Hierarchical Federated Learning Framework for Semantic Communication in 6G LEO Satellites
- Trans4Trans: Efficient Transformer for Transparent Object Segmentation to Help Visually Impaired People Navigate in the Real World
- CAMELTrack: Context-Aware Multi-cue ExpLoitation for Online Multi-Object Tracking
- Improving Routing in Sparse Mixture of Experts with Graph of Tokens
- Towards a Foundation-Model Paradigm for Aerodynamic Prediction in Three-dimensional Design
- FOVI: A biologically-inspired foveated interface for deep vision models
- M2MRF: Many-to-Many Reassembly of Features for Tiny Lesion Segmentation in Fundus Images
- Modeling Scientific Experiment Scenes: Dataset and Model
- GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection
- FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers
- Higher-Order Fourier Neural Operator: Explicit Mode Mixer for Nonlinear PDEs
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
- Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
- SemGS: Feed-Forward Semantic 3D Gaussian Splatting from Sparse Views for Generalizable Scene Understanding
- Pack-PTQ: Advancing Post-training Quantization of Neural Networks by Pack-wise Reconstruction
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
- Automatic detection of fin, operculum and skin deformities in Mediterranean Fish Species
- Vision Transformers Need More Than Registers
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
- Flowers: A Warp Drive for Neural PDE Solvers
- NPMixer: Hierarchical Neighboring Patch Mixing for Time Series Forecasting
- Vision Transformers in Precision Agriculture: A Comprehensive Survey
- Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces
- ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
- Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions
- A Test Suite for Efficient Robustness Evaluation of Face Recognition Systems
- XeMap: Contextual Referring in Large-Scale Remote Sensing Environments
- Comparison of Different Deep Neural Network Models in the Cultural Heritage Domain
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
- LDPoly: Latent Diffusion for Polygonal Road Outline Extraction in Large-Scale Topographic Mapping
- SNR-aware Semantic Image Transmission with Deep Learning-based Channel Estimation in Fading Channels
- LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs
- SCOPE-MRI: Bankart Lesion Detection as a Case Study in Data Curation and Deep Learning for Challenging Diagnoses
- Leveraging Depth Maps and Attention Mechanisms for Enhanced Image Inpainting
- SteelBlastQC: Shot-blasted Steel Surface Dataset with Interpretable Detection of Surface Defects
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
- Enhancing breast cancer detection on screening mammogram using self-supervised learning and a hybrid deep model of Swin Transformer and Convolutional Neural Network
- Multi-axis Analysis of Image Manipulation Localization
- Hybrid Quantum-MambaVision: A Quantum-enhanced State Space Model for Calibrated Mixed-type Wafer Defect Detection
- Comparison of Image Processing Models in Quark Gluon Jet Classification
- Breast Cancer Detection from Multi-View Screening Mammograms with Visual Prompt Tuning
- Reinforcement Learning-Based Heterogeneous Multi-Task Optimization in Semantic Broadcast Communications
- BARIS: Boundary-Aware Refinement with Environmental Degradation Priors for Robust Underwater Instance Segmentation
- BiXiao: An AI-Based Atmospheric Environment Forecasting Model Using Discontinuous Grids
- GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
- ODExAI: A Comprehensive Object Detection Explainable AI Evaluation
- Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
- Global Climate Model Bias Correction Using Deep Learning
- MLICv2: Enhanced Multi-Reference Entropy Modeling for Learned Image Compression
- Application of a Mixture of Experts-based Foundation Model to the GlueX DIRC Detector
- CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- AIBuildAI: An AI Agent for Automatically Building AI Models
- Dream-Box: Object-wise Outlier Generation for Out-of-Distribution Detection
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain Structures
- Toward Accurate and Reliable Iris Segmentation Using Uncertainty Learning
- Group Downsampling with Equivariant Anti-aliasing
- 3D Deep-learning-based Segmentation of Human Skin Sweat Glands and Their 3D Morphological Response to Temperature Variations
- A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
- DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
- Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities
- Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
- From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
- STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
- A multilevel approach to accelerate the training of Transformers
- SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
- ResT: An Efficient Transformer for Visual Recognition
- Mamba-Sea: A Mamba-based Framework with Global-to-Local Sequence Augmentation for Generalizable Medical Image Segmentation
- CKMDiff: A Generative Diffusion Model for CKM Construction via Inverse Problems with Learned Priors
- Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
- Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
- Geometry-informed multimodal variational autoencoder for real-time prediction of properties for Ti–6Al–4V fabricated using PBF-LB
- Time-adaptive Video Frame Interpolation based on Residual Diffusion
- Content-Distortion High-Order Interaction for Blind Image Quality Assessment
- Transformer representation learning is necessary for dynamic multi-modal physiological data on small-cohort patients
- EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively
- RouteWinFormer: A Route-Window Transformer for Middle-range Attention in Image Restoration
- Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
- Almost Right: Making First-Layer Kernels Nearly Orthogonal Improves Model Generalization
- Systematic Literature Review on Vehicular Collaborative Perception -- A Computer Vision Perspective
- A multi-scale vision transformer-based multimodal GeoAI model for mapping Arctic permafrost thaw
- Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
- Iterative Collaboration Network Guided By Reconstruction Prior for Medical Image Super-Resolution
- MTSGL: Multi-Task Structure Guided Learning for Robust and Interpretable SAR Aircraft Recognition
- Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
- MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
- Opening the black box of deep learning: Validating the statistical association between explainable artificial intelligence (XAI) and clinical domain knowledge in fundus image-based glaucoma diagnosis
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
- Integrating Non-Linear Radon Transformation for Diabetic Retinopathy Grading
- Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
- Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
- SuoiAI: Building a Dataset for Aquatic Invertebrates in Vietnam
- Dynamic 3D KAN Convolution with Adaptive Grid Optimization for Hyperspectral Image Classification
- NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study
- ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- Hybrid Knowledge Transfer through Attention and Logit Distillation for On-Device Vision Systems in Agricultural IoT
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- MSAD-Net: Multiscale and Spatial Attention-based Dense Network for Lung Cancer Classification
- NTIRE 2025 Challenge on Image Super-Resolution (×4): Methods and Results
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
- Rethinking Traffic Flow Forecasting: From Transition to Generatation
- Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
- Fighting Fires from Space: Leveraging Vision Transformers for Enhanced Wildfire Detection and Characterization
- Towards Accurate and Interpretable Neuroblastoma Diagnosis via Contrastive Multi-scale Pathological Image Analysis
- DAM-Net: Domain Adaptation Network with Micro-Labeled Fine-Tuning for Change Detection
- DenSe-AdViT: A novel Vision Transformer for Dense SAR Object Detection
- ViG3D-UNet: Volumetric Vascular Connectivity-Aware Segmentation via 3D Vision Graph Representation
- FocusNet: Transformer-enhanced Polyp Segmentation with Local and Pooling Attention
- Depth-aware RGB-D concrete crack segmentation and quantification using progressive cross-modal attention
- A Novel Hybrid Approach for Retinal Vessel Segmentation with Dynamic Long-Range Dependency and Multi-Scale Retinal Edge Fusion Enhancement
- Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
- HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
- Multiscale Tensor Summation Factorization as a New Neural Network Layer (MTS Layer) for Multidimensional Data Processing
- VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture
- SAR Object Detection with Self-Supervised Pretraining and Curriculum-Aware Sampling
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
- Region-Specific Evaluation of Plaque Segmentation in Cross-sectional Projections of Carotid Ultrasound Images Using Deep Learning Models in a Sub-clinical Atherosclerosis Cohort
- Efficient Masked Image Compression with Position-Indexed Self-Attention
- Expert Kernel Generation Network Driven by Contextual Mapping for Hyperspectral Image Classification
- Hierarchical Feature Learning for Medical Point Clouds via State Space Model
- NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results
- zkVC: Fast Zero-Knowledge Proof for Private and Verifiable Computing
- A Review of YOLOv12: Attention-Based Enhancements vs. Previous Versions
- A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning
- Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
- DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
- Search is All You Need for Few-shot Anomaly Detection
- Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
- URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation
- Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
- APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
- Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI
- Invisible Shortcuts: Why Vision Encoders Know Your Camera
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
- Ge2mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
- DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Graph Network for Sign Language Tasks
- Mamba-Based Ensemble learning for White Blood Cell Classification
- DDFusion:Degradation-Decoupled Fusion Framework for Robust Infrared and Visible Images Fusion
- 3D Wavelet Convolutions with Extended Receptive Fields for Hyperspectral Image Classification
- ConvShareViT: Enhancing Vision Transformers with Convolutional Attention Mechanisms for Free-Space Optical Accelerators
- Defending Against Frequency-Based Attacks with Diffusion Models
- Making Acoustic Side-Channel Attacks on Noisy Keyboards Viable with LLM-Assisted Spectrograms' "Typo" Correction
- Lightweight Medical Image Restoration via Integrating Reliable Lesion-Semantic Driven Prior
- Enhanced Small Target Detection via Multi-Modal Fusion and Attention Mechanisms: A YOLOv5 Approach
- Noise2Ghost: Self-supervised deep convolutional reconstruction for ghost imaging
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
- GFT: Gradient Focal Transformer
- Global and Local Mamba Network for Multi-Modality Medical Image Super-Resolution
- IGL-DT: Iterative Global-Local Feature Learning with Dual-Teacher Semantic Segmentation Framework under Limited Annotation Scheme
- DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
- Semantic Depth Matters: Explaining Errors of Deep Vision Networks through Perceived Class Similarities
- UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
- Low-Light Image Enhancement using Event-Based Illumination Estimation
- Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
- SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- D2iT: Dynamic Diffusion Transformer for Accurate Image Generation
- ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance
- A CNN-based Local-Global Self-Attention via Averaged Window Embeddings for Hierarchical ECG Analysis
- Multi-Modal Brain Tumor Segmentation via 3D Multi-Scale Self-attention and Cross-attention
- Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition
- AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
- Evolved Hierarchical Masking for Self-Supervised Learning
- Mixture of Group Experts for Learning Invariant Representations
- Exploring Synergistic Ensemble Learning: Uniting CNNs, MLP-Mixers, and Vision Transformers to Enhance Image Classification
- Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
- A Hybrid Fully Convolutional CNN-Transformer Model for Inherently Interpretable Disease Detection from Retinal Fundus Images
- On Background Bias of Post-Hoc Concept Embeddings in Computer Vision DNNs
- Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- Adversarial Examples in Environment Perception for Automated Driving (Review)
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
- On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
- Benchmarking Image Embeddings for E-Commerce: Evaluating Off-the Shelf Foundation Models, Fine-Tuning Strategies and Practical Trade-offs
- SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion
- Novel Pooling-based VGG-Lite for Pneumonia and Covid-19 Detection from Imbalanced Chest X-Ray Datasets
- ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method
- Learning Optimal Prompt Ensemble for Multi-source Visual Prompt Transfer
- Crafting Query-Aware Selective Attention for Single Image Super-Resolution
- Distilling Textual Priors from LLM to Efficient Image Fusion
- Optuna vs Code Llama: Are LLMs a New Paradigm for Hyperparameter Tuning?
- DefMamba: Deformable Visual State Space Model
- Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
- D-Feat Occlusions: Diffusion Features for Robustness to Partial Visual Occlusions in Object Recognition
- A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
- KunPeng: A Global Ocean Environmental Model
- Feature Importance-Aware Deep Joint Source-Channel Coding for Computationally Efficient and Adjustable Image Transmission
- DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
- DSwinIR: Rethinking Window-based Attention for Image Restoration
- Synthetic frequency patterns injection for data-agnostic deepfake detection
- MIMRS: A Survey on Masked Image Modeling in Remote Sensing
- SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
- Optimizing Specific and Shared Parameters for Efficient Parameter Tuning
- Multi-encoder nnU-Net outperforms transformer models with self-supervised pretraining
- NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices
- Charm: The Missing Piece in ViT fine-tuning for Image Aesthetic Assessment
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Computer vision‐based real‐time cable safety assessment under vehicle‐induced bridge fires
- Beyond Conventional Transformers: The Medical X-ray Attention (MXA) Block for Improved Multi-Label Diagnosis Using Knowledge Distillation
- APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
- Chinese Paper-Cutting Style Transfer via Vision Transformer
- Adaptive Frequency Enhancement Network for Remote Sensing Image Semantic Segmentation
- HGFormer: Topology-Aware Vision Transformer with HyperGraph Learning
Discussions
Related