Emerging Properties in Self-Supervised Vision Transformers
2021/04/29 by Caron, Mathilde, Touvron, Hugo, Misra, Ishan +4 · 781 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2104.14294
Abstract
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder, multi-crop training, and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
Cited by
- Contour Information Aware 2D Gaussian Splatting for Image Representation
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion
- Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
- 3D Scene Change Modeling With Consistent Multi-View Aggregation
- Improved cystic hygroma detection from prenatal imaging using ultrasound-specific self-supervised representation learning
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Normalizing Trajectory Models
- A satellite foundation model for improved wealth monitoring
- Semantic Semi-Incremental Data-Association-Free Object SLAM
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- OrganLens: Organ-Specific Representation Learning for CT Foundation Models
- Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
- DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
- Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
- Robustifying pathology foundation models via fine-tuning
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models
- Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
- GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
- Zero-Shot Segmentation through Prototype-Guidance for Multi-Label Plant Species Identification
- Vehicle-centric Perception via Multimodal Structured Pre-training
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors
- Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
- Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
- WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
- Multi-Part Object Representations via Graph Structures and Co-Part Discovery
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- Next-Embedding Prediction Makes Strong Vision Learners
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- DeContext as Defense: Safe Image Editing in Diffusion Transformers
- In Pursuit of Pixel Supervision for Visual Pre-training
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
- Borrowing from anything: A generalizable framework for reference-guided instance editing
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- Unified Semantic Transformer for 3D Scene Understanding
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- Recurrent Video Masked Autoencoders
- Do-Undo: Generating and Reversing Physical Actions in Vision-Language Models
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
- Calibrating Uncertainty for Zero-Shot Adversarial CLIP
- Sharpness-aware Dynamic Anchor Selection for Generalized Category Discovery
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- SPDMark: Selective Parameter Displacement for Robust Video Watermarking
- V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
- RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- ClusIR: Towards Cluster-Guided All-in-One Image Restoration
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
- VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
- StainNet: Scaling Self-Supervised Foundation Models on Immunohistochemistry and Special Stains for Computational Pathology
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- Composing Concepts from Images and Videos via Concept-prompt Binding
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- Self-Supervised Learning with Gaussian Processes
- CytoDINO: Risk-Aware and Biologically-Informed Adaptation of DINOv3 for Bone Marrow Cytomorphology
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- Advancing Autonomous Driving System Testing: Demands, Challenges, and Future Directions
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Is Generation Required for Data-Efficient Perception?
- Relational Visual Similarity
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- Unified Video Editing with Temporal Reasoner
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- Generalized Geometry Encoding Volume for Real-time Stereo Matching
- Transferring Clinical Knowledge into ECGs Representation
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Label-Efficient Point Cloud Segmentation with Active Learning
- InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
- Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm
- EvoIR: Towards All-in-One Image Restoration via Evolutionary Frequency Modulation
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
- ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Unsupervised Structural Scene Decomposition via Foreground-Aware Slot Attention with Pseudo-Mask Guidance
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- Language-Guided Open-World Anomaly Segmentation
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation
- From Regression to Classification: Exploring the Benefits of Categorical Representations of Energy in MLIPs
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Multilingual Training-Free Remote Sensing Image Captioning
- TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- Learning to Predict Aboveground Biomass from RGB Images with 3D Synthetic Scenes
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- Adversarial Flow Models
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
- Video Generation Models Are Good Latent Reward Models
- MeanFlow Transformers with Representation Autoencoders
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- DINO-Tok: Adapting DINO for Visual Tokenizers
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
- Automated Histopathologic Assessment of Hirschsprung Disease Using a Multi-Stage Vision Transformer Framework
- A Training-Free Approach for Multi-ID Customization via Attention Adjustment and Spatial Control
- VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild
- FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation
- Rethinking Semi-Supervised Node Classification with Self-Supervised Graph Clustering
- Annotation-Free Class-Incremental Learning
- RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- stable-pretraining-v1: Foundation Model Research Made Simple
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Point-to-Point: Sparse Motion Guidance for Controllable Video Editing
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
- Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
- TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
- Graph Neural Networks for Surgical Scene Segmentation
- Exploiting Inter-Sample Information for Long-tailed Out-of-Distribution Detection
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
- PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
- From Low-Rank Features to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
- A Dataset and Baseline for Deep Learning-Based Visual Quality Inspection in Remanufacturing
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
- Training-free Detection of AI-generated images via Cropping Robustness
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Tissue Aware Nuclei Detection and Classification Model for Histopathology Images
- MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
- Passive Dementia Screening via Facial Temporal Micro-Dynamics Analysis of In-the-Wild Talking-Head Video
- DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning
- Extremal Contours: Gradient-driven contours for compact visual attribution
- Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
- CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
- Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
- EgoCogNav: Cognition-aware Human Egocentric Navigation
- Φeat: Physically Grounded Material Feature Representation
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- Fast Data Attribution for Text-to-Image Models
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
- Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
- SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- TomoGraphView: 3D Medical Image Classification with Omnidirectional Slice Representations and Graph Neural Networks
- SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
- Expanding the Content-Style Frontier: a Balanced Subspace Blending Approach for Content-Style LoRA Fusion
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- Mitigating Negative Flips via Margin Preserving Training
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- WEDepth: Efficient Adaptation of World Knowledge for Monocular Depth Estimation
- Exploring the Underwater World Segmentation without Extra Training
- Visual Bridge: Universal Visual Perception Representations Generating
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- Detecting Generated Images by Fitting Natural Image Distributions
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- Local K-Similarity Constraint for Federated Learning with Label Noise
- CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
- Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory
- Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
- Finetuning-Free Personalization of Text to Image Generation via Hypernetworks
- Accelerating Physical Property Reasoning for Augmented Visual Cognition
- An Augmentation Overlap Theory of Contrastive Learning
- Web-Scale Collection of Video Data for 4D Animal Reconstruction
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks
- PLUTO-4: Frontier Pathology Foundation Models
- Dynamic Reflections: Probing Video Representations with Text Alignment
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- Self-Supervised Moving Object Segmentation of Sparse and Noisy Radar Point Clouds
- Differentiable Hierarchical Visual Tokenization
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
- Generalized Category Discovery under Domain Shift: A Frequency Domain Perspective
- Challenging DINOv3 Foundation Model under Low Inter-Class Variability: A Case Study on Fetal Brain Ultrasound
- MIFO: Learning and Synthesizing Multi-Instance from One Image
- Who Made This? Fake Detection and Source Attribution with Diffusion Features
- CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
- Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
- A filtering scheme for confocal laser endomicroscopy (CLE)-video sequences for self-supervised learning
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- ROGR: Relightable 3D Objects using Generative Relighting
- Image Quality Dependent Degradation for AI Systems
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image Completion
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion
- Is Dimensionality a Barrier for Retrieval Models?
- Zero-shot World Models Are Developmentally Efficient Learners
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- RL makes MLLMs see better than SFT
- Generating metamers of human scene understanding
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- A multimodal whole-slide foundation model for pathology
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Eigenfunction Extraction for Ordered Representation Learning
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- A Unified Geometric Space Bridging AI Models and the Human Brain
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- Neural USD: An object-centric framework for iterative editing and control
- LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Reliable Robotic Task Execution in the Face of Anomalies
- Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
- FastJAM: a Fast Joint Alignment Model for Images
- MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- FrameShield: Adversarially Robust Video Anomaly Detection
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Weak-to-Strong Generalization under Distribution Shifts
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Model Merging with Functional Dual Anchors
- VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization
- What Does It Take to Build a Performant Selective Classifier?
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Exploring Conditions for Diffusion models in Robotic Control
- I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- Brain-Inspired Perspective on Configurations: Unsupervised Similarity and Early Cognition
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- CBDiff:Conditional Bernoulli Diffusion Models for Image Forgery Localization
- Transformed Multi-view 3D Shape Features with Contrastive Learning
- 3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
- An Explainable Hybrid AI Framework for Enhanced Tuberculosis and Symptom Detection
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- SPACeR: Self-Play Anchoring with Centralized Reference Models
- Elastic ViTs from Pretrained Models without Retraining
- Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
- DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
- Closed-Loop Transfer for Weakly-supervised Affordance Grounding
- HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery
- 2D3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
- Learning After Model Deployment
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- NeuCo-Bench: A Novel Benchmark Framework for Neural Embeddings in Earth Observation
- Region in Context: Text-condition Image editing with Human-like semantic reasoning
- Unsupervised Monocular Road Segmentation for Autonomous Driving via Scene Geometry
- Universal and Transferable Attacks on Pathology Foundation Models
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Salient Concept-Aware Generative Data Augmentation
- Towards Adversarial Robustness and Uncertainty Quantification in DINOv2-based Few-Shot Anomaly Detection
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Synchronization of Multiple Videos
- Scaling Vision Transformers for Functional MRI with Flat Maps
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion
- Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
- Universal Image Restoration Pre-training via Masked Degradation Classification
- OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- AnyUp: Universal Feature Upsampling
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Diffusion Transformers with Representation Autoencoders
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- G2L:From Giga-Scale to Cancer-Specific Large-Scale Pathology Foundation Models via Knowledge Distillation
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Data or Language Supervision: What Makes CLIP Better than DINO?
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Uncovering Semantic Selectivity of Latent Groups in Higher Visual Cortex with Mutual Information-Guided Diffusion
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- An uncertainty-aware framework for data-efficient multi-view animal pose estimation
- Cell Instance Segmentation: The Devil Is in the Boundaries
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- TARO: Toward Semantically Rich Open-World Object Detection
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- Vision Language Models: A Survey of 26K Papers
- LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Sparse components distinguish visual pathways & their alignment to neural networks
- Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- DADO: A Depth-Attention framework for Object Discovery
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
- Heptapod: Language Modeling on Visual Signals
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Teleportraits: Training-Free People Insertion into Any Scene
- Human3R: Everyone Everywhere All at Once
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Visual Representations inside the Language Model
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention Disentanglement
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- Knowledge Distillation Detection for Open-weights Models
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- ActiveMark: on watermarking of visual foundation models via massive activations
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- Hierarchical Generalized Category Discovery for Brain Tumor Classification in Digital Pathology
- Self-Supervised Representation Learning as Mutual Information Maximization
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Feature Identification for Hierarchical Contrastive Learning
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Rehearsal-free and Task-free Online Continual Learning With Contrastive Prompt
- Can World Models Benefit VLMs for World Dynamics?
- From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
- Cutting the Skip: Training Residual-Free Transformers
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- The Impact of Scaling Training Data on Adversarial Robustness
- Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- TTT3R: 3D Reconstruction as Test-Time Training
- "Where Can I Park?" Understanding Human Perspectives and Scalably Detecting Disability Parking from Aerial Imagery
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- HBSplat: Robust Sparse-View Gaussian Reconstruction with Hybrid-Loss Guided Depth and Bidirectional Warping
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Semantic Editing with Coupled Stochastic Differential Equations
- Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- Generalized Category Discovery in Hyperspectral Images via Prototype Subspace Modeling
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
- Griffin: Generative Reference and Layout Guided Image Composition
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- Multi-Modal Manipulation via Multi-Modal Policy Consensus
- MindCraft: How Concept Trees Take Shape In Deep Models
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation
- PartCo: Part-Level Correspondence Priors Enhance Category Discovery
- Orochi: Versatile Biomedical Image Processor
- Category Discovery: An Open-World Perspective
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- On the Status of Foundation Models for SAR Imagery
- Large Material Gaussian Model for Relightable 3D Generation
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- Dense Semantic Matching with VGGT Prior
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive Modeling
- CusEnhancer: A Zero-Shot Scene and Controllability Enhancement Method for Photo Customization via ResInversion
- An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
- Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
- Anatomically Constrained Transformers for Cardiac Amyloidosis Classification
- Enhancing Transformer-Based Vision Models: Addressing Feature Map Anomalies Through Novel Optimization Strategies
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging
- 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
- Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- OSF: On Pre-training and Scaling of Sleep Foundation Models
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- MAPO: Mixed Advantage Policy Optimization
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Training-Free Multi-Style Fusion Through Reference-Based Adaptive Modulation
- Global Minimizers of Sigmoid Contrastive Loss
- Prompt-DAS: Annotation-Efficient Prompt Learning for Domain Adaptive Semantic Segmentation of Electron Microscopy Images
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Can multimodal representation learning by alignment preserve modality-specific information?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Overview of PlantCLEF 2022: Image-based plant identification at global scale
- M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- Animalbooth: multimodal feature enhancement for animal subject personalization
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training
- Causal Fingerprints of AI Generative Models
- Explaining deep learning for ECG using time-localized clusters
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- Communication Efficient Split Learning of ViTs with Attention-based Double Compression
- DinoTwins: Combining DINO and Barlow Twins for Robust, Label-Efficient Vision Transformers
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Data Leakage in Visual Datasets
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- SAMIR, an efficient registration framework via robust feature learning from SAM
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- Superpixel Anything: A general object-based framework for accurate yet regular superpixel segmentation
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- Data Scaling Laws for Radiology Foundation Models
- Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
- Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
- DRAG: Data Reconstruction Attack using Guided Diffusion
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
- When marine radar target detection meets pretrained large language models
- Domain-Adaptive Pretraining Improves Primate Behavior Recognition
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- ToMA: Token Merge with Attention for Diffusion Models
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- WebSight: A Vision-First Architecture for Robust Web Agents
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
- MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- EfficientIML: Efficient High-Resolution Image Manipulation Localization
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- RINO: Renormalization Group Invariance with No Labels
- Understanding Ice Crystal Habit Diversity with Self-Supervised Learning
- Reconstruction Alignment Improves Unified Multimodal Models
- CellEcoNet: Decoding the Cellular Language of Pathology with Deep Learning for Invasive Lung Adenocarcinoma Recurrence Prediction
- Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations
- QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients
- Video-based Generalized Category Discovery via Memory-Guided Consistency-Aware Contrastive Learning
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Prior Distribution and Model Confidence
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
- Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
- Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
- Towards Open World Detection: A Survey
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- Weakly-Supervised Learning of Dense Functional Correspondences
- Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)
- Unsupervised Instance Segmentation with Superpixels
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- RiverScope: High-Resolution River Masking Dataset
- Vision encoders should be image size agnostic and task driven
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- SALAD -- Semantics-Aware Logical Anomaly Detection
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- MedDINOv3: How to adapt vision foundation models for medical image segmentation?
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
- PractiLight: Practical Light Control Using Foundational Diffusion Models
- IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects
- Decomposing and Revising What Language Models Generate
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
- Standardized Multi-Layer Tissue Maps for Enhanced Artificial Intelligence Integration and Search in Large-Scale Whole Slide Image Archives
- SatDINO: A Deep Dive into Self-Supervised Pretraining for Remote Sensing
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Representation Learning with Adaptive Superpixel Coding
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Beyond Imaging: Vision Transformer Digital Twin Surrogates for 3D+T Biological Tissue Dynamics
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- Self-supervised structured object representation learning
- Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
- SCAR: A Characterization Scheme for Multi-Modal Dataset
- DNP-Guided Contrastive Reconstruction with a Reverse Distillation Transformer for Medical Anomaly Detection
- Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion
- Efficient Multi-Source Knowledge Transfer by Model Merging
- Articulate3D: Zero-Shot Text-Driven 3D Object Posing
- Style4D-Bench: A Benchmark Suite for 4D Stylization
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- Dual-Distilled Heterogeneous Federated Learning with Adaptive Margins for Trainable Global Prototypes
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- Controllable Latent Space Augmentation for Digital Pathology
- LookOut: Real-World Humanoid Egocentric Navigation
- OmniTry: Virtual Try-On Anything without Masks
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- Unleashing Semantic and Geometric Priors for 3D Scene Completion
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Structure-preserving Feature Alignment for Old Photo Colorization
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- Contrastive Representations for Temporal Reasoning
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- Leveraging Diffusion Models for Stylization using Multiple Style Images
- Splat Feature Solver
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
- SPG: Style-Prompting Guidance for Style-Specific Content Creation
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Remove360: Benchmarking Residuals After Object Removal in 3D Gaussian Splatting
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Better Supervised Fine-tuning for VQA: Integer-Only Loss
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Generalizable Federated Learning using Client Adaptive Focal Modulation
- Self-Supervised Stereo Matching with Multi-Baseline Contrastive Learning
- VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
- Separating Knowledge and Perception with Procedural Data
- Per-Query Visual Concept Learning
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
- CObL: Toward Zero-Shot Ordinal Layering without User Prompting
- X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
- Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a User-Centric Optimization Space
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Symmetry Understanding of 3D Shapes via Chirality Disentanglement
- SMOL-MapSeg: Show Me One Label as prompt
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
- Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- AdvDINO: Domain-Adversarial Self-Supervised Representation Learning for Spatial Proteomics
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- No Masks Needed: Explainable AI for Deriving Segmentation from Classification
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- Learning Robust Intervention Representations with Delta Embeddings
- Benchmarking Foundation Models for Mitotic Figure Classification
- SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- FFHQ-Makeup: Paired Synthetic Makeup Dataset with Facial Consistency Across Multiple Styles
- MILD: Multi-Layer Diffusion Strategy for Complex and Precise Multi-IP Aware Human Erasing
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Elucidating the Role of Feature Normalization in IJEPA
- Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal Retrieval
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- Subject or Style: Adaptive and Training-Free Mixture of LoRAs
- Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Enhancing Object Discovery for Unsupervised Instance Segmentation and Object Detection
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
- AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and Editing
- A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
- Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles
- IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Steering Guidance for Personalized Text-to-Image Diffusion Models
- Towards Robust Semantic Correspondence: A Benchmark and Insights
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Training-free Geometric Image Editing on Diffusion Models
- iLRM: An Iterative Large 3D Reconstruction Model
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- FaceGCD: Generalized Face Discovery via Dynamic Prefix Generation
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Unlocking Interpretability for RF Sensing: A Complex-Valued White-Box Transformer
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- BANG: Dividing 3D Assets via Generative Exploded Dynamics
- Ensemble Foreground Management for Unsupervised Object Discovery
- Compositional Video Synthesis by Temporal Object-Centric Learning
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- Harnessing Diffusion-Yielded Score Priors for Image Restoration
Related