Emerging Properties in Self-Supervised Vision Transformers
2021/04/29 by Mathilde Caron, Caron, Mathilde, Hugo Touvron +12 · 1,489 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CV
paper · pdf · doi:10.48550/arxiv.2104.14294
21 pages
arxiv created 2021/05/24 · arxiv updated 2021/05/25
Abstract
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder, multi-crop training, and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
Citations
Cited by
- Contour Information Aware 2D Gaussian Splatting for Image Representation
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion
- Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
- 3D Scene Change Modeling With Consistent Multi-View Aggregation
- Improved cystic hygroma detection from prenatal imaging using ultrasound-specific self-supervised representation learning
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Normalizing Trajectory Models
- A satellite foundation model for improved wealth monitoring
- Semantic Semi-Incremental Data-Association-Free Object SLAM
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- OrganLens: Organ-Specific Representation Learning for CT Foundation Models
- Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
- DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
- Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
- Robustifying pathology foundation models via fine-tuning
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models
- Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
- GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
- Zero-Shot Segmentation through Prototype-Guidance for Multi-Label Plant Species Identification
- Vehicle-centric Perception via Multimodal Structured Pre-training
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors
- Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
- Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
- WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
- Multi-Part Object Representations via Graph Structures and Co-Part Discovery
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- Next-Embedding Prediction Makes Strong Vision Learners
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- DeContext as Defense: Safe Image Editing in Diffusion Transformers
- In Pursuit of Pixel Supervision for Visual Pre-training
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
- Borrowing from anything: A generalizable framework for reference-guided instance editing
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- Unified Semantic Transformer for 3D Scene Understanding
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- Recurrent Video Masked Autoencoders
- Do-Undo Bench: Reversibility for Action Understanding in Image Generation
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
- Calibrating Uncertainty for Zero-Shot Adversarial CLIP
- Sharpness-aware Dynamic Anchor Selection for Generalized Category Discovery
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- SPDMark: Selective Parameter Displacement for Robust Video Watermarking
- V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
- RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- ClusIR: Towards Cluster-Guided All-in-One Image Restoration
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
- VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
- StainNet: Scaling Self-Supervised Foundation Models on Immunohistochemistry and Special Stains for Computational Pathology
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- Composing Concepts from Images and Videos via Concept-prompt Binding
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- Self-Supervised Learning with Gaussian Processes
- CytoDINO: Risk-Aware and Biologically-Informed Adaptation of DINOv3 for Bone Marrow Cytomorphology
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- Advancing Autonomous Driving System Testing: Demands, Challenges, and Future Directions
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Is Generation Required for Data-Efficient Perception?
- Relational Visual Similarity
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- VideoCoF: Unified Video Editing with Temporal Reasoner
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- Generalized Geometry Encoding Volume for Real-time Stereo Matching
- Transferring Clinical Knowledge into ECGs Representation
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Label-Efficient Point Cloud Segmentation with Active Learning
- InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
- Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm
- EvoIR: Towards All-in-One Image Restoration via Evolutionary Frequency Modulation
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
- ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Unsupervised Structural Scene Decomposition via Foreground-Aware Slot Attention with Pseudo-Mask Guidance
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- Language-Guided Open-World Anomaly Segmentation
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation
- From Regression to Classification: Exploring the Benefits of Categorical Representations of Energy in MLIPs
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Multilingual Training-Free Remote Sensing Image Captioning
- TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- Learning to Predict Aboveground Biomass from RGB Images with 3D Synthetic Scenes
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- Adversarial Flow Models
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency
- Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
- Video Generation Models Are Good Latent Reward Models
- MeanFlow Transformers with Representation Autoencoders
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- DINO-Tok: Adapting DINO for Visual Tokenizers
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
- Automated Histopathologic Assessment of Hirschsprung Disease Using a Multi-Stage Vision Transformer Framework
- A Training-Free Approach for Multi-ID Customization via Attention Adjustment and Spatial Control
- VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild
- FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation
- Rethinking Semi-Supervised Node Classification with Self-Supervised Graph Clustering
- Annotation-Free Class-Incremental Learning
- RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- stable-pretraining-v1: Foundation Model Research Made Simple
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Point-to-Point: Sparse Motion Guidance for Controllable Video Editing
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
- Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
- TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
- Graph Neural Networks for Surgical Scene Segmentation
- Exploiting Inter-Sample Information for Long-tailed Out-of-Distribution Detection
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
- PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
- From Low-Rank Features to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
- A Dataset and Baseline for Deep Learning-Based Visual Quality Inspection in Remanufacturing
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
- Training-free Detection of AI-generated images via Cropping Robustness
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Tissue Aware Nuclei Detection and Classification Model for Histopathology Images
- MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
- Passive Dementia Screening via Facial Temporal Micro-Dynamics Analysis of In-the-Wild Talking-Head Video
- DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning
- Extremal Contours: Gradient-driven contours for compact visual attribution
- Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
- CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
- Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
- EgoCogNav: Cognition-aware Human Egocentric Navigation
- Φeat: Physically Grounded Material Feature Representation
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- Fast Data Attribution for Text-to-Image Models
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
- Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
- SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- TomoGraphView: 3D Medical Image Classification with Omnidirectional Slice Representations and Graph Neural Networks
- SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
- Expanding the Content-Style Frontier: a Balanced Subspace Blending Approach for Content-Style LoRA Fusion
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- Mitigating Negative Flips via Margin Preserving Training
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- WEDepth: Efficient Adaptation of World Knowledge for Monocular Depth Estimation
- Exploring the Underwater World Segmentation without Extra Training
- Visual Bridge: Universal Visual Perception Representations Generating
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- Detecting Generated Images by Fitting Natural Image Distributions
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- Local K-Similarity Constraint for Federated Learning with Label Noise
- CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
- Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory
- Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
- Finetuning-Free Personalization of Text to Image Generation via Hypernetworks
- Accelerating Physical Property Reasoning for Augmented Visual Cognition
- An Augmentation Overlap Theory of Contrastive Learning
- Web-Scale Collection of Video Data for 4D Animal Reconstruction
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks
- PLUTO-4: Frontier Pathology Foundation Models
- Dynamic Reflections: Probing Video Representations with Text Alignment
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- Self-Supervised Moving Object Segmentation of Sparse and Noisy Radar Point Clouds
- Differentiable Hierarchical Visual Tokenization
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
- Generalized Category Discovery under Domain Shift: A Frequency Domain Perspective
- Challenging DINOv3 Foundation Model under Low Inter-Class Variability: A Case Study on Fetal Brain Ultrasound
- MIFO: Learning and Synthesizing Multi-Instance from One Image
- Who Made This? Fake Detection and Source Attribution with Diffusion Features
- CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
- Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
- A filtering scheme for confocal laser endomicroscopy (CLE)-video sequences for self-supervised learning
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- ROGR: Relightable 3D Objects using Generative Relighting
- Image Quality Dependent Degradation for AI Systems
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Improving Out-of-Distribution Detection via Dynamic Covariance Calibration
- GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image Completion
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion
- Is Dimensionality a Barrier for Retrieval Models?
- Zero-shot World Models Are Developmentally Efficient Learners
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- RL makes MLLMs see better than SFT
- Generating metamers of human scene understanding
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- A multimodal whole-slide foundation model for pathology
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Eigenfunction Extraction for Ordered Representation Learning
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- A Unified Geometric Space Bridging AI Models and the Human Brain
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- Neural USD: An object-centric framework for iterative editing and control
- LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Reliable Robotic Task Execution in the Face of Anomalies
- Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
- FastJAM: a Fast Joint Alignment Model for Images
- MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- FrameShield: Adversarially Robust Video Anomaly Detection
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Weak-to-Strong Generalization under Distribution Shifts
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Model Merging with Functional Dual Anchors
- VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- Adversarial Concept Distillation for One-Step Diffusion Personalization
- What Does It Take to Build a Performant Selective Classifier?
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Exploring Conditions for Diffusion models in Robotic Control
- I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
- A theoretical framework for self-supervised contrastive learning for continuous dependent data
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- Brain-Inspired Perspective on Configurations: Unsupervised Similarity and Early Cognition
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- CBDiff:Conditional Bernoulli Diffusion Models for Image Forgery Localization
- Transformed Multi-view 3D Shape Features with Contrastive Learning
- 3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
- An Explainable Hybrid AI Framework for Enhanced Tuberculosis and Symptom Detection
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- SPACeR: Self-Play Anchoring with Centralized Reference Models
- Elastic ViTs from Pretrained Models without Retraining
- Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
- DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
- Closed-Loop Transfer for Weakly-supervised Affordance Grounding
- HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery
- 2D3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
- Learning After Model Deployment
- MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- NeuCo-Bench: A Novel Benchmark Framework for Neural Embeddings in Earth Observation
- Region in Context: Text-condition Image editing with Human-like semantic reasoning
- Unsupervised Monocular Road Segmentation for Autonomous Driving via Scene Geometry
- Universal and Transferable Attacks on Pathology Foundation Models
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Salient Concept-Aware Generative Data Augmentation
- Towards Adversarial Robustness and Uncertainty Quantification in DINOv2-based Few-Shot Anomaly Detection
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Synchronization of Multiple Videos
- Scaling Vision Transformers for Functional MRI with Flat Maps
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion
- Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
- Universal Image Restoration Pre-training via Masked Degradation Classification
- OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Only-Style: Stylistic Consistency in Image Generation without Content Leakage
- AnyUp: Universal Feature Upsampling
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Diffusion Transformers with Representation Autoencoders
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- G2L:From Giga-Scale to Cancer-Specific Large-Scale Pathology Foundation Models via Knowledge Distillation
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Data or Language Supervision: What Makes CLIP Better than DINO?
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Uncovering Semantic Selectivity of Latent Groups in Higher Visual Cortex with Mutual Information-Guided Diffusion
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- An uncertainty-aware framework for data-efficient multi-view animal pose estimation
- Cell Instance Segmentation: The Devil Is in the Boundaries
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- TARO: Toward Semantically Rich Open-World Object Detection
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- Vision Language Models: A Survey of 26K Papers
- LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Sparse components distinguish visual pathways & their alignment to neural networks
- Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- DADO: A Depth-Attention framework for Object Discovery
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
- Heptapod: Language Modeling on Visual Signals
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Teleportraits: Training-Free People Insertion into Any Scene
- Human3R: Everyone Everywhere All at Once
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Visual Representations inside the Language Model
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention Disentanglement
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- Knowledge Distillation Detection for Open-weights Models
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- ActiveMark: on watermarking of visual foundation models via massive activations
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- Hierarchical Generalized Category Discovery for Brain Tumor Classification in Digital Pathology
- Self-Supervised Representation Learning as Mutual Information Maximization
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Feature Identification for Hierarchical Contrastive Learning
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Rehearsal-free and Task-free Online Continual Learning With Contrastive Prompt
- Can World Models Benefit VLMs for World Dynamics?
- From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
- Cutting the Skip: Training Residual-Free Transformers
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- The Impact of Scaling Training Data on Adversarial Robustness
- Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- TTT3R: 3D Reconstruction as Test-Time Training
- "Where Can I Park?" Understanding Human Perspectives and Scalably Detecting Disability Parking from Aerial Imagery
- Benchmarking ECG FMs: A Reality Check Across Clinical Tasks
- CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- HBSplat: Robust Sparse-View Gaussian Reconstruction with Hybrid-Loss Guided Depth and Bidirectional Warping
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Semantic Editing with Coupled Stochastic Differential Equations
- Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- Generalized Category Discovery in Hyperspectral Images via Prototype Subspace Modeling
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
- Griffin: Generative Reference and Layout Guided Image Composition
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- Multi-Modal Manipulation via Multi-Modal Policy Consensus
- MindCraft: How Concept Trees Take Shape In Deep Models
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation
- PartCo: Part-Level Correspondence Priors Enhance Category Discovery
- Orochi: Versatile Biomedical Image Processor
- Category Discovery: An Open-World Perspective
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- On the Status of Foundation Models for SAR Imagery
- Large Material Gaussian Model for Relightable 3D Generation
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- Dense Semantic Matching with VGGT Prior
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive Modeling
- CusEnhancer: A Zero-Shot Scene and Controllability Enhancement Method for Photo Customization via ResInversion
- An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
- Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
- Anatomically Constrained Transformers for Cardiac Amyloidosis Classification
- Enhancing Transformer-Based Vision Models: Addressing Feature Map Anomalies Through Novel Optimization Strategies
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging
- 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
- Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- OSF: On Pre-training and Scaling of Sleep Foundation Models
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- MAPO: Mixed Advantage Policy Optimization
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Training-Free Multi-Style Fusion Through Reference-Based Adaptive Modulation
- Global Minimizers of Sigmoid Contrastive Loss
- Prompt-DAS: Annotation-Efficient Prompt Learning for Domain Adaptive Semantic Segmentation of Electron Microscopy Images
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Can multimodal representation learning by alignment preserve modality-specific information?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Overview of PlantCLEF 2022: Image-based plant identification at global scale
- M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- Animalbooth: multimodal feature enhancement for animal subject personalization
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training
- Causal Fingerprints of AI Generative Models
- Explaining deep learning for ECG using time-localized clusters
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- Communication Efficient Split Learning of ViTs with Attention-based Double Compression
- DinoTwins: Combining DINO and Barlow Twins for Robust, Label-Efficient Vision Transformers
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Data Leakage in Visual Datasets
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- SAMIR, an efficient registration framework via robust feature learning from SAM
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- Superpixel Anything: A general object-based framework for accurate yet regular superpixel segmentation
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- Data Scaling Laws for Radiology Foundation Models
- Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
- Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
- DRAG: Data Reconstruction Attack using Guided Diffusion
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
- When marine radar target detection meets pretrained large language models
- Domain-Adaptive Pretraining Improves Primate Behavior Recognition
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- ToMA: Token Merge with Attention for Diffusion Models
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- WebSight: A Vision-First Architecture for Robust Web Agents
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
- MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- EfficientIML: Efficient High-Resolution Image Manipulation Localization
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- RINO: Renormalization Group Invariance with No Labels
- Understanding Ice Crystal Habit Diversity with Self-Supervised Learning
- Reconstruction Alignment Improves Unified Multimodal Models
- CellEcoNet: Decoding the Cellular Language of Pathology with Deep Learning for Invasive Lung Adenocarcinoma Recurrence Prediction
- Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations
- QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients
- Video-based Generalized Category Discovery via Memory-Guided Consistency-Aware Contrastive Learning
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- WakeupUrban: Unsupervised Semantic Segmentation of Mid-20th century Urban Landscapes with Satellite Imagery
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Prior Distribution and Model Confidence
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
- Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
- Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
- Towards Open World Detection: A Survey
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- Weakly-Supervised Learning of Dense Functional Correspondences
- Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)
- Unsupervised Instance Segmentation with Superpixels
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- RiverScope: High-Resolution River Masking Dataset
- Vision encoders should be image size agnostic and task driven
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- SALAD -- Semantics-Aware Logical Anomaly Detection
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- MedDINOv3: How to adapt vision foundation models for medical image segmentation?
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
- PractiLight: Practical Light Control Using Foundational Diffusion Models
- IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects
- Decomposing and Revising What Language Models Generate
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
- Standardized Multi-Layer Tissue Maps for Enhanced Artificial Intelligence Integration and Search in Large-Scale Whole Slide Image Archives
- SatDINO: A Deep Dive into Self-Supervised Pretraining for Remote Sensing
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Representation Learning with Adaptive Superpixel Coding
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Beyond Imaging: Vision Transformer Digital Twin Surrogates for 3D+T Biological Tissue Dynamics
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- Self-supervised structured object representation learning
- Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
- SCAR: A Characterization Scheme for Multi-Modal Dataset
- DNP-Guided Contrastive Reconstruction with a Reverse Distillation Transformer for Medical Anomaly Detection
- Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion
- Efficient Multi-Source Knowledge Transfer by Model Merging
- Articulate3D: Zero-Shot Text-Driven 3D Object Posing
- Style4D-Bench: A Benchmark Suite for 4D Stylization
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- Dual-Distilled Heterogeneous Federated Learning with Adaptive Margins for Trainable Global Prototypes
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- Controllable Latent Space Augmentation for Digital Pathology
- LookOut: Real-World Humanoid Egocentric Navigation
- OmniTry: Virtual Try-On Anything without Masks
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- Unleashing Semantic and Geometric Priors for 3D Scene Completion
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning
- Structure-preserving Feature Alignment for Old Photo Colorization
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- Contrastive Representations for Temporal Reasoning
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- Leveraging Diffusion Models for Stylization using Multiple Style Images
- Splat Feature Solver
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
- SPG: Style-Prompting Guidance for Style-Specific Content Creation
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Remove360: Benchmarking Residuals After Object Removal in 3D Gaussian Splatting
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Better Supervised Fine-tuning for VQA: Integer-Only Loss
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Generalizable Federated Learning using Client Adaptive Focal Modulation
- Self-Supervised Stereo Matching with Multi-Baseline Contrastive Learning
- VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
- Separating Knowledge and Perception with Procedural Data
- Per-Query Visual Concept Learning
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
- CObL: Toward Zero-Shot Ordinal Layering without User Prompting
- X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
- Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a User-Centric Optimization Space
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Symmetry Understanding of 3D Shapes via Chirality Disentanglement
- SMOL-MapSeg: Show Me One Label as prompt
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
- Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- AdvDINO: Domain-Adversarial Self-Supervised Representation Learning for Spatial Proteomics
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- No Masks Needed: Explainable AI for Deriving Segmentation from Classification
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- Learning Robust Intervention Representations with Delta Embeddings
- Benchmarking Foundation Models for Mitotic Figure Classification
- SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- FFHQ-Makeup: Paired Synthetic Makeup Dataset with Facial Consistency Across Multiple Styles
- MILD: Multi-Layer Diffusion Strategy for Complex and Precise Multi-IP Aware Human Erasing
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Elucidating the Role of Feature Normalization in IJEPA
- Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal Retrieval
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- Subject or Style: Adaptive and Training-Free Mixture of LoRAs
- Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Enhancing Object Discovery for Unsupervised Instance Segmentation and Object Detection
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
- AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and Editing
- A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
- Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles
- Discovering and using Spelke segments
- IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Steering Guidance for Personalized Text-to-Image Diffusion Models
- Towards Robust Semantic Correspondence: A Benchmark and Insights
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Training-free Geometric Image Editing on Diffusion Models
- iLRM: An Iterative Large 3D Reconstruction Model
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- FaceGCD: Generalized Face Discovery via Dynamic Prefix Generation
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Towards White-Box Deep Wireless Sensing
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- BANG: Dividing 3D Assets via Generative Exploded Dynamics
- ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
- Ensemble Foreground Management for Unsupervised Object Discovery
- Compositional Video Synthesis by Temporal Object-Centric Learning
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- Harnessing Diffusion-Yielded Score Priors for Image Restoration
- JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
- HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
- Improving Joint Embedding Predictive Architecture with Diffusion Noise
- Latent Denoising Makes Good Tokenizers
- Diffusion Beats Autoregressive in Data-Constrained Settings
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
- Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
- ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
- HQ-SMem: Video Segmentation and Tracking Using Memory Efficient Object Embedding With Selective Update and Self-Supervised Distillation Feedback
- Patch Pruning Strategy Based on Robust Statistical Measures of Attention Weight Diversity in Vision Transformers
- Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
- Improving Personalized Image Generation through Social Context Feedback
- Identifying Prompted Artist Names from Generated Images
- Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
- InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
- eMargin: Revisiting Contrastive Learning with Margin-Based Separation
- OPRD: On-Policy Representation Distillation
- Image Generators are Generalist Vision Learners
- PRAGMA: Revolut Foundation Model
- Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection
- DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF
- Unified Latents (UL): How to train your latents
- On the robustness of modeling grounded word learning through a child's egocentric input
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted Attention
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
- PositionIC: Unified Position and Identity Consistency for Image Customization
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
- Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
- Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
- Attention (as Discrete-Time Markov) Chains
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
- Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
- A High Magnifications Histopathology Image Dataset for Oral Squamous Cell Carcinoma Diagnosis and Prognosis
- Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
- DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
- UniLGL: Learning Uniform Place Recognition for FOV-limited/Panoramic LiDAR Global Localization
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
- PASS: Peer-agreement based sample selection for training with instance dependent noisy labels
- International multicenter validation of AI-driven ultrasound detection of ovarian cancer
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- Grounded language acquisition through the eyes and ears of a single child
- BERT Bi-modal self-supervised learning for crop classification using Sentinel-2 and Planetscope
- Adapting a global plant identification model to detect invasive alien plant species in high-resolution road side images
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- FADE: A Task-Agnostic Upsampling Operator for Encoder–Decoder Architectures
- DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Microscopy Images
- State of the Art on Diffusion Models for Visual Computing
- Semi-Supervised Classification and Segmentation on High Resolution Aerial Images
- SimMIM: A Simple Framework for Masked Image Modeling
- On Masked Pre-training and the Marginal Likelihood
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- Deep Learning with Label Differential Privacy
- Jointly Learnable Data Augmentations for Self-Supervised GNNs
- Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
- Global Interaction Modelling in Vision Transformer via Super Tokens
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Vector-quantized Image Modeling with Improved VQGAN
- Offline Clustering Approach to Self-supervised Learning for Class-imbalanced Image Data
- A Simple Long-Tailed Recognition Baseline via Vision-Language Model
- BEiT: BERT Pre-Training of Image Transformers
- Efficient Self-supervised Vision Transformers for Representation Learning
- DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- ResNet strikes back: An improved training procedure in timm
- Exploring the Limits of Out-of-Distribution Detection
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Revisiting Model Stitching to Compare Neural Representations
- IA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Is it Time to Replace CNNs with Transformers for Medical Images?
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Algorithm Fairness in AI for Medicine and Healthcare
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- XCiT: Cross-Covariance Image Transformers
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Intriguing Properties of Vision Transformers
- Self-Distilled Self-Supervised Representation Learning
- Towards Autonomous Riding: A Review of Perception, Planning, and Control in Intelligent Two-Wheelers
- SpatialTrackerV2: 3D Point Tracking Made Easy
- Cluster Contrast for Unsupervised Visual Representation Learning
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation
- Vision Transformer for Learning Driving Policies in Complex Multi-Agent Environments
- From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
- Automated cell annotation and classification on histopathology for spatial biomarker discovery
- MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis
- RS-MTDF: Multi-Teacher Distillation and Fusion for Remote Sensing Semi-Supervised Semantic Segmentation
- MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
- Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs
- UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
- MMOne: Representing Multiple Modalities in One Scene
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
- Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
- Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
- CodeBrain: Bridging Decoupled Tokenizer and Multi-Scale Architecture for EEG Foundation Model
- SensorLM: Learning the Language of Wearable Sensors
- Vision Language Action Models in Robotic Manipulation: A Systematic Review
- 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- Reprogramming Vision Foundation Models for Spatio-Temporal Forecasting
- TolerantECG: A Foundation Model for Imperfect Electrocardiogram
- Foundation Models in Medical Imaging: A Review and Outlook
- From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation
- Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
- Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies
- Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- Fine-Grained Spatially Varying Material Selection in Images
- Self-supervised pretraining of vision transformers for animal behavioral analysis and neural encoding
- CKAA: Cross-subspace Knowledge Alignment and Aggregation for Robust Continual Learning
- Intention-Conditioned Flow Occupancy Models
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Stable Score Distillation
- Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
- Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
- Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
- MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification
- Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
- Single-pass Adaptive Image Tokenization for Minimum Program Search
- MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- A Principled Framework for Multi-View Contrastive Learning
- Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic Segmentation
- Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
- EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision
- Divergence-Based Similarity Function for Multi-View Contrastive Learning
- FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning
- LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
- LOVON: Legged Open-Vocabulary Object Navigator
- Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
- Does Data Scaling Lead to Visual Compositional Generalization?
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
- DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
- T-LoRA: Single Image Diffusion Model Customization Without Overfitting
- SingLoRA: Low Rank Adaptation Using a Single Matrix
- MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
- Generative Panoramic Image Stitching
- Semi-weakly Supervised Contrastive Representation Learning for Retinal Fundus Images
- Geometric-Guided Few-Shot Dental Landmark Detection with Human-Centric Foundation Model
- QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation
- Critique of World Model
- Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
- Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
- MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
- Clustering via Self-Supervised Diffusion
- Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
- Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
- Zero-shot Inexact CAD Model Alignment from a Single Image
- FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
- No time to train! Training-Free Reference-Based Instance Segmentation
- RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
- Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning
- Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
- PosDiffAE: Position-aware Diffusion Auto-encoder For High-Resolution Brain Tissue Classification Incorporating Artifact Restoration
- Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
- A computationally frugal open-source foundation model for thoracic disease detection in lung cancer screening programs
- Robust brain age estimation from structural MRI with contrastive learning
- FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
- Component Adaptive Clustering for Generalized Category Discovery
- ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
- AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
- SPoT: Subpixel Placement of Tokens in Vision Transformers
- NOCTIS: Novel Object Cyclic Threshold based Instance Segmentation
- Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
- TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
- Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
- DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting
- Few-shot Classification as Multi-instance Verification: Effective Backbone-agnostic Transfer across Domains
- VISTA: Open-Vocabulary, Task-Relevant Robot Exploration with Online Semantic Gaussian Splatting
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- Do Protein Transformers Have Biological Intelligence?
- Reducing Variability of Multiple Instance Learning Methods for Digital Pathology
- FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
- Emergent musical properties of a transformer under contrastive self-supervised learning
- Pruning by Block Benefit: Exploring the Properties of Vision Transformer Blocks during Domain Adaptation
- When Test-Time Adaptation Meets Self-Supervised Models
- Immune organization in sentinel lymph nodes of melanoma patients is prognostic of distant metastases
- UFMTrack, an Under-Flow Migration Tracker enabling analysis of the entire multi-step immune cell extravasation cascade across the blood-brain barrier in microfluidic devices
- Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
- GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data
- FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation
- High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
- Dare to Plagiarize? Plagiarized Painting Recognition and Retrieval
- Self-Supervised Contrastive Learning for Multi-Label Images
- VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point Clouds
- Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
- SAM2Auto: Auto Annotation Using FLASH
- Identifiable Object Representations under Spatial Ambiguities
- Embodied AI Agents: Modeling the World
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
- SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
- Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
- HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
- Weakly Supervised Object Segmentation by Background Conditional Divergence
- Enhanced Consistency Bi-directional GAN(CBiGAN) for Malware Anomaly Detection
- InvZW: Invariant Feature Learning via Noise-Adversarial Training for Robust Image Zero-Watermarking
- DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
- Multiple Object Stitching for Unsupervised Representation Learning
- From 2D to 3D Cognition: A Brief Survey of General World Models
- Vision Transformers Don't Need Trained Registers
- RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
- Disentangled representations of microscopy images
- General Methods Make Great Domain-specific Foundation Models: A Case-study on Fetal Ultrasound
- Distillation-Enabled Knowledge Alignment for Generative Semantic Communications in AIGC Provisioning Tasks
- Stylized Structural Patterns for Improved Neural Network Pre-training
- Is an object-centric representation beneficial for robotic manipulation ?
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Image Reconstruction as a Tool for Feature Analysis
- Auto-Regressively Generating Multi-View Consistent Images
- ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
- DIP: Unsupervised Dense In-Context Post-training of Visual Representations
- Selective Social-Interaction via Individual Importance for Fast Human Trajectory Prediction
- Improving Black-Box Generative Attacks via Generator Semantic Consistency
- Benchmarking Foundation Models and Parameter-Efficient Fine-Tuning for Prognosis Prediction in Medical Imaging
- Limitations of NERF with pre-trained Vision Features for Few-Shot 3D Reconstruction
- Joint Embedding Predictive Architecture for self-supervised pretraining on polymer molecular graphs
- Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing
- BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
- Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual Invariance
- Emergent Temporal Correspondences from Video Diffusion Transformers
- Class Agnostic Instance-level Descriptor for Visual Instance Search
- Few-Shot Generalized Category Discovery With Retrieval-Guided Decision Boundary Enhancement
- Bridging Brain with Foundation Models through Self-Supervised Learning
- Evaluating the fairness of fine-tuning strategies in self-supervised learning
- Reliable Few-shot Learning under Dual Noises
- Deep Learning Foundation Models from Classical Molecular Descriptors
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- Interpretability and Generalization Bounds for Learning Spatial Physics
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- MapFM: Foundation Model-Driven HD Mapping with Multi-Task Contextual Learning
- Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot Manipulation
- Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation
- Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection
- Discrete JEPA: Learning Discrete Token Representations without Reconstruction
- Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
- Latent Action Diffusion for Cross-Embodiment Manipulation
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
- Contrastive Self-Supervised Learning As Neural Manifold Packing
- Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
- TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast
- Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention
- Evolution of ReID: From Early Methods to LLM Integration
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Long-Tailed Learning for Generalized Category Discovery
- Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation
- Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
- Generalized Category Discovery under the Long-Tailed Distribution
- EKPC: Elastic Knowledge Preservation and Compensation for Class-Incremental Learning
- EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
- Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
- UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning
- MRI-CORE: A Foundation Model for Magnetic Resonance Imaging
- Visual Pre-Training on Unlabeled Images using Reinforcement Learning
- Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology
- Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
- Demonstrating Multi-Suction Item Picking at Scale via Multi-Modal Learning of Pick Success
- Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
- ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- Accurate and efficient zero-shot 6D pose estimation with frozen foundation models
- Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks
- Improving Wildlife Out-of-Distribution Detection: Africas Big Five
- Splat and Replace: 3D Reconstruction with Repetitive Elements
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
- Spatially-Enhanced Recurrent Memory for Long-Range Mapless Navigation via End-to-End Reinforcement Learning
- Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
- Reliable Evaluation of MRI Motion Correction: Dataset and Insights
- Contrastive Flow Matching
- Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels
- OGGSplat: Open Gaussian Growing for Generalizable Reconstruction with Expanded Field-of-View
- Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
- PixCell: A generative foundation model for digital histopathology images
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- FLAM: Frame-Wise Language-Audio Modeling
- FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
- TD-TOG Dataset: Benchmarking Zero-Shot and One-Shot Task-Oriented Grasping for Object Generalization
- FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
- Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
- Line of Sight: On Linear Representations in VLLMs
- LSM-2: Learning from Incomplete Wearable Sensor Data
- Towards Reliable Identification of Diffusion-based Image Manipulations
- Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
- The mutual exclusivity bias of bilingual visually grounded speech models
- Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging
- Characterizing Structural Regularities of Labeled Data in Overparameterized Models
- Learning from Similarity Proportion Loss for Classifying Skeletal Muscle Recovery Stages
- RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
- Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- Transformer-Based Source-Free Domain Adaptation
- Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
- SemiOccam: A Robust Semi-Supervised Image Recognition Network Using Sparse Labels
- Object-level Self-Distillation for Vision Pretraining
- Image Editing As Programs with Diffusion Models
- Learning Smooth State-Dependent Traversability from Dense Point Clouds
- How PARTs assemble into wholes: Learning the relative composition of images
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- Attention-Only Transformers via Unrolled Subspace Denoising
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- PartComposer: Learning and Composing Part-Level Concepts from Single-Image Examples
- RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS
- Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
- Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
- Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
- Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images
- Guiding Registration with Emergent Similarity from Pre-Trained Diffusion Models
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
- Attacking Attention of Foundation Models Disrupts Downstream Tasks
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- Talk2SAM: Text-Guided Semantic Enhancement for Complex-Shaped Object Segmentation
- SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
- Auto-Annotation with Expert-Crafted Guidelines: A Study through 3D LiDAR Detection Benchmark
- Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning
- FORLA: Federated Object-centric Representation Learning with Slot Attention
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning
- G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
- SEMNAV: A Semantic Segmentation-Driven Approach to Visual Semantic Navigation
- On Improving Adversarial Transferability of Vision Transformers
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
- E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
- MoCA: Multi-modal Cross-masked Autoencoder for Time Series in Digital Health
- SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
- Sparse Imagination for Efficient Visual World Model Planning
- Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity
- Generating Synthetic Data via Augmentations for Improved Facial Resemblance in DreamBooth and InstantID
- ECP-Mamba: An Efficient Multi-scale Self-supervised Contrastive Learning Method with State Space Model for PolSAR Image Classification
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
- Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
- A versatile foundation model for cine cardiac magnetic resonance image analysis tasks
- Video Signature: Implicit Watermarking for Video Diffusion Models
- SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
- SST: Self-training with Self-adaptive Thresholding for Semi-supervised Learning
- Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
- A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Tackling View-Dependent Semantics in 3D Language Gaussian Splatting
- KairosAD: A SAM-Based Model for Industrial Anomaly Detection on Embedded Devices
- Benchmarking Foundation Models for Zero-Shot Biometric Tasks
- Understanding while Exploring: Semantics-driven Active Mapping
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
- Representational Difference Explanations
- TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
- Skin Lesion Phenotyping via Nested Multi-modal Contrastive Learning
- Mobi-π: Mobilizing Your Robot Learning Policy
- Scalable Class-Centric Visual Interactive Labeling
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- Federated Unsupervised Semantic Segmentation
- SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
- ICEv2: Interpretability, Comprehensiveness, and Explainability in Vision Transformer
- Anomalies by Synthesis: Anomaly Detection using Generative Diffusion Models for Off-Road Navigation
- AquaMonitor: A multimodal multi-view image sequence dataset for real-life aquatic invertebrate biodiversity monitoring
- SineLoRAΔ: Sine-Activated Delta Compression
- SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- A Survey on Training-free Open-Vocabulary Semantic Segmentation
- Defining Foundation Models for Computational Science: A Call for Clarity and Rigor
- Long-Short Temporal Contrastive Learning of Video Transformers
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- Object Concepts Emerge from Motion
- ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
- Vision Transformers with Self-Distilled Registers
- Multi-instance Learning as Downstream Task of Self-Supervised Learning-based Pre-trained Model
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
- The Missing Point in Vision Transformers for Universal Image Segmentation
- Benchmarking Detection Transfer Learning with Vision Transformers
- A Contrastive Learning Foundation Model Based on Perfectly Aligned Sample Pairs for Remote Sensing Images
- Exploring the Possibility of TypiClust for Low-Budget Federated Active Learning
- Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
- SeisCoDE: 3D Seismic Interpretation Foundation Model with Contrastive Self-Distillation Learning
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
- DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
- DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
- Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
- AmorLIP: Efficient Language-Image Pretraining via Amortization
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
- Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
- C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging
- Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
- ETHead: Generating Expressive 3D Facial Animation and Head Movement from Speech
- Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
- Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling
- Test-Time Scaling of Diffusion Models via Noise Trajectory Search
- Generative AI and foundation models in medical image
- Self-Organizing Visual Prototypes for Non-Parametric Representation Learning
- REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
- Attention Mechanisms in Computer Vision: A Survey
- Semantic Correspondence: Unified Benchmarking and a Strong Baseline
- SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
- ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
- Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
- A Coreset Selection of Coreset Selection Literature: Introduction and Recent Advances
- From Flight to Insight: Semantic 3D Reconstruction for Aerial Inspection via Gaussian Splatting and Language-Guided Segmentation
- Learning Shared Representations from Unpaired Data
- Transformer brain encoders explain human high-level visual responses
- Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
- Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data
- HOFT: Householder Orthogonal Fine-tuning
- Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
- T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
- The P3 dataset: Pixels, Points and Polygons for Multimodal Building Vectorization
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification
- Lung Nodule-SSM: Self-Supervised Lung Nodule Detection and Classification in Thoracic CT Images
- VET-DINO: Learning Anatomical Understanding Through Multi-View Distillation in Veterinary Imaging
- gen2seg: Generative Models Enable Generalizable Instance Segmentation
- Generative AI for Autonomous Driving: A Review
- Can machines learn to see without visual databases?
- Emerging Properties in Unified Multimodal Pretraining
- SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification
- SSR: Similarity-Shift Refinement for Training-Free Object-Centric Masks
- Generalized Category Discovery via Token Manifold Capacity Learning
- SuperMapNet for Long-Range and High-Accuracy Vectorized HD Map Construction
- Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
- Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
- Collaborative Unlabeled Data Optimization
- RECON: Robust symmetry discovery via Explicit Canonical Orientation Normalization
- FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
- GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
- Semi-Supervised Vision Transformers
- RGB-to-Polarization Estimation: A New Task and Benchmark Study
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
- Simplicity is Key: An Unsupervised Pretraining Approach for Sparse Radio Channels
- Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
- Industrial Synthetic Segment Pre-training
- AdaDim: Dimensionality Adaptation for SSL Representational Dynamics
- Training Latent Diffusion Models with Interacting Particle Algorithms
- Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
- The perceptual primacy of feeling: Affectless visual machines explain a majority of variance in human visually evoked affect
- Transformer-based unsupervised contrastive learning for histopathological image classification
- No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
- Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
- PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
- Task-agnostic Low-rank Residual Adaptation for Efficient Federated Continual Fine-Tuning
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
- MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
- Equally Critical: Samples, Targets, and Their Mappings in Datasets
- Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer
- Continuous Subspace Optimization for Continual Learning
- iSegMan: Interactive Segment-and-Manipulate 3D Gaussians
- Are vision language models robust to uncertain inputs?
- Cross-Model Transfer of Task Vectors via Few-Shot Orthogonal Alignment
- Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models
- CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
- CleanPatrick: A Benchmark for Image Data Cleaning
- Is a knife the same as a plunger? Comparing the attentional effects of weapons and non-threatening unusual objects in dynamic scenes
- A densely sampled and richly annotated acoustic data set from a wild bird population
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- MutualNeRF: Improve the Performance of NeRF under Limited Samples with Mutual Information Theory
- Object-Centric Representations Improve Policy Generalization in Robot Manipulation
- Self-supervised perception for tactile skin covered dexterous hands
- The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders
- CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
- Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
- A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
- GAIA: A Foundation Model for Operational Atmospheric Dynamics
- EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
- IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation
- Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers
- CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
- Don't Forget your Inverse DDIM for Image Editing
- Air-Ground Collaboration for Language-Specified Missions in Unknown Environments
- Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
- Few-shot Novel Category Discovery
- Distilling Drifting Transformers with Representation Autoencoders
- Vision Foundation Model Embedding-Based Semantic Anomaly Detection
- Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
- A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
- Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
- Image Classification Using a Diffusion Model as a Pre-Training Model
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
- Towards Embodiment Scaling Laws in Robot Locomotion
- CGTrack: Cascade Gating Network with Hierarchical Feature Aggregation for UAV Tracking
- Towards a Unified Representation Evaluation Framework Beyond Downstream Tasks
- Split Matching for Inductive Zero-shot Semantic Segmentation
- Do Vision Transformers See Like Convolutional Neural Networks?
- PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding
- Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning
- Lifelong Whole Slide Image Analysis: Online Vision-Language Adaptation and Past-to-Present Gradient Distillation
- Video Generation Models are General-Purpose Vision Learners
- VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
- CostFilter-AD: Enhancing Anomaly Detection through Matching Cost Filtering
- Self-Supervision Enhances Instance-based Multiple Instance Learning Methods in Digital Pathology: A Benchmark Study
- Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
- Diffusion-based Adversarial Purification from the Perspective of the Frequency Domain
- DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
- Self-Supervised Learning by Estimating Twin Class Distributions
- Taming Outlier Tokens in Diffusion Transformers
- Teaching an Agent to Sketch One Part at a Time
- SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE
- The Role of Artificial Intelligence in the SKA Era
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- INSID3: Training-Free In-Context Segmentation with DINOv3
- KVT: k-NN Attention for Boosting Vision Transformers
- VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
- Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
- Estimating the Robustness of Classification Models by the Structure of the Learned Feature-Space
- InstructAttribute: Fine-grained Object Attributes editing with Instruction
- Online Federation For Mixtures of Proprietary Agents with Black-Box Encoders
- Vision Transformers Need More Than Registers
- Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
- Unpaired Image-to-Image Translation via a Self-Supervised Semantic Bridge
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- On the Dynamics of Observation and Semantics
- Learn from your own latents and not from tokens: A sample-complexity theory
- When Does LeJEPA Learn a World Model?
- Segmentation-Guided Spatial Indexing for Generalizable and Explainable Deepfake Detection
- Recursive KL Divergence Optimization: A Dynamic Framework for Representation Learning
- OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
- PaintCopilot: Modeling Painting as Autonomous Artistic Continuation
- Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space
- STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
- Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
- NearID: Identity Representation Learning via Near-identity Distractors
- PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations
- In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
- DeepAndes: A Self-Supervised Vision Foundation Model for Multi-Spectral Remote Sensing Imagery of the Andes
- CompleteMe: Reference-based Human Image Completion
- Enhancing breast cancer detection on screening mammogram using self-supervised learning and a hybrid deep model of Swin Transformer and Convolutional Neural Network
- Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
- UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- Self-supervised local learning rules learn the hidden hierarchical structure of high-dimensional data
- MARCO: Navigating the Unseen Space of Semantic Correspondence
- The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery
- MUFASA: A Multi-Layer Framework for Slot Attention
- Generative Visual Code Mobile World Models
- Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
- LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
- Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning
- MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis
- OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
- Frequency-Forcing: From Scaling-as-Time to Soft Frequency Guidance
- KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains
- CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
- Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- SemanticMoments: Training-Free Motion Similarity via Third Moment Features
- DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
- Flow-based Extremal Mathematical Structure Discovery
- OceanSAR-2: A Universal Feature Extractor for SAR Ocean Observation
- The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
- RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
- Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
- Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
- Speaker Verification Under Real Classroom Conditions for English Speech
- From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
- Improving Open-World Object Localization by Discovering Background
- Enhancing rice breeding efficiency through semi-supervised detection and segmentation of panicles and leaves
- SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
- Fine-tune Smarter, Not Harder: Parameter-Efficient Fine-Tuning for Geospatial Foundation Models
- RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
- A Genealogy of Foundation Models in Remote Sensing
- Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
- NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
- Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision
- ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
- SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
- Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
- Seeking Physics in Diffusion Noise
- Training state-of-the-art pathology foundation models with orders of magnitude less data
- CountingDINO: A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones
- Facial Foundational Model Advances Early Warning of Coronary Artery Disease from Live Videos with DigitalShadow
- Credal Self-Supervised Learning
- DreamO: A Unified Framework for Image Customization
- I-Con: A Unifying Framework for Representation Learning
- ForesightNav: Learning Scene Imagination for Efficient Exploration
- Active Learning with a Noisy Annotator
- CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss
- Representation Synthesis by Probabilistic Many-Valued Logic Operation in Self-Supervised Learning
- Boosting Generative Image Modeling via Joint Image-Feature Synthesis
- Pose Optimization for Autonomous Driving Datasets using Neural Rendering Models
- UINO-FSS: Unifying Representation Learning and Few-shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelation
- Automated Measurement of Eczema Severity with Self-Supervised Learning
- "I Know It When I See It": Mood Spaces for Connecting and Expressing Visual Concepts
- HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
- LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
- Points2Polygons: Context-Based Segmentation from Weak Labels Using Adversarial Networks
- Image Editing with Diffusion Models: A Survey
- Interpretable Single-View 3D Gaussian Splatting using Unsupervised Hierarchical Disentangled Representation Learning
- Personalized Text-to-Image Generation with Auto-Regressive Models
- Digital Twin Generation from Visual Data: A Survey
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
- GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
- A Survey of Pathology Foundation Model: Progress and Future Directions
- CAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
- AdaVid: Adaptive Video-Language Pretraining
- Search is All You Need for Few-shot Anomaly Detection
- JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
- SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation
- MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
- CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
- Invisible Shortcuts: Why Vision Encoders Know Your Camera
- EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding
- StyleComposer: Training-Free Multi-Reference Style Composition
- Position: It's Time to Optimize LLMs for Self-Consistency
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
- Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
- ViMo: A Generative Visual GUI World Model for App Agents
- FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- Pay Attention to What and Where? Interpretable Feature Extractor in Vision-based Deep Reinforcement Learning
- Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
- An Image is Worth K Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
- GFT: Gradient Focal Transformer
- Efficient Generative Model Training via Embedded Representation Warmup
- Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
- On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
- Causal integration of chemical structures improves representations of microscopy images for morphological profiling
- Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
- Evolved Hierarchical Masking for Self-Supervised Learning
- Flux Already Knows -- Activating Subject-Driven Image Generation without Training
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- Steering CLIP's vision transformer with sparse autoencoders
- Boosting multi-demographic federated learning for chest radiograph analysis using general-purpose self-supervised representations
- Enhancing knowledge retention for continual learning with domain-specific adapters and features gating
- Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
- FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
- Detect Anything 3D in the Wild
- MARS: a Multimodal Alignment and Ranking System for Few-Shot Segmentation
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- Revisiting Likelihood-Based Out-of-Distribution Detection by Modeling Representations
- Benchmarking Image Embeddings for E-Commerce: Evaluating Off-the Shelf Foundation Models, Fine-Tuning Strategies and Practical Trade-offs
- WS-DETR: Robust Water Surface Object Detection through Vision-Radar Fusion with Detection Transformer
- Fast Adaptation with Behavioral Foundation Models
- Self-Bootstrapping for Versatile Test-Time Adaptation
- GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces
- Impact of Language Guidance: A Reproducibility Study
- Efficient Self-Supervised Learning for Earth Observation via Dynamic Dataset Curation
- Latent Diffusion U-Net Representations Contain Positional Embeddings and Anomalies
- Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
- EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
- Are We Done with Object-Centric Learning?
- Measuring Déjà vu Memorization Efficiently
- econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians
- InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
- MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
- Hyperbolic Category Discovery
- Deep-learning–based assessment of postoperative nasolabial morphology in unilateral cleft lip repair
- Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
- PEAKS: Selecting Key Training Examples Incrementally via Prediction Error Anchored by Kernel Similarity
- Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
- Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
- DebGCD: Debiased Learning with Distribution Guidance for Generalized Category Discovery
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
- Conditional Object-Centric Learning from Video
- REJEPA: A Novel Joint-Embedding Predictive Architecture for Efficient Remote Sensing Image Retrieval
Related