Masked Autoencoders Are Scalable Vision Learners
2021/11/11 by Kaiming He, Xinlei Chen, He, Kaiming +9 · 4 voices · 776 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.2111.06377
arxiv published 2021/11/11 · arxiv updated 2021/12/19
Abstract
This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3x or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pre-training and shows promising scaling behavior.
Citations
Cited by
- Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
- Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers
- Improved cystic hygroma detection from prenatal imaging using ultrasound-specific self-supervised representation learning
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- A Scale-adaptive Vision Model Links C. elegans Neuronal Morphology to Behavior for Neurotoxicity Assessment
- Self-Supervised Learning from Noisy and Incomplete Data
- Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
- Multiclass Classification without Labels via Posterior Simplex Geometry
- Beam-Response Contrastive Learning for Transmitter-Side MIMO CSI Representation
- How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models
- Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data
- Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
- Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
- SLIM-Brain: A Data- and Training-Efficient Foundation Model for fMRI Data Analysis
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- BertsWin: Resolving Topological Sparsity in 3D Masked Autoencoders via Component-Balanced Structural Optimization
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Surgical Scene Segmentation using a Spike-Driven Video Transformer with Real-Time Potential
- SparScene: Efficient Traffic Scene Representation via Sparse Graph Learning for Large-Scale Trajectory Generation
- Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition
- Granular-ball Guided Masking: Structure-aware Data Augmentation
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- FGDCC: Fine-Grained Deep Cluster Categorization -- A Framework for Intra-Class Variability Problems in Plant Classification
- Block-Recurrent Dynamics in Vision Transformers
- SemanticGen: Video Generation in Semantic Space
- Vehicle-centric Perception via Multimodal Structured Pre-training
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- Multi-Modal Soccer Scene Analysis with Masked Pre-Training
- InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Phase-space entropy at acquisition reflects downstream learnability
- A Study of Finetuning Video Transformers for Multi-view Geometry Tasks
- brat: Aligned Multi-View Embeddings for Brain MRI Analysis
- MCVI-SANet: A lightweight semi-supervised model for LAI and SPAD estimation of winter wheat under vegetation index saturation
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection
- Towards Pixel-Wise Anomaly Location for High-Resolution PCBA via Self-Supervised Image Reconstruction
- Next-Embedding Prediction Makes Strong Vision Learners
- SARMAE: Masked Autoencoder for SAR Representation Learning
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- In Pursuit of Pixel Supervision for Visual Pre-training
- Topological Metric for Unsupervised Embedding Quality Evaluation
- Keep the Core: Adversarial Priors for Significance-Preserving Brain MRI Segmentation
- Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
- From Feature Interaction to Feature Generation: A Generative Paradigm of CTR Prediction Models
- LCMem: A Universal Model for Robust Image Memorization Detection
- SuperCLIP: CLIP with Simple Classification Supervision
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- Recurrent Video Masked Autoencoders
- Self-Supervised Ultrasound Representation Learning for Renal Anomaly Prediction in Prenatal Imaging
- RecTok: Reconstruction Distillation along Rectified Flow
- SkyCap: Bitemporal VHR Optical-SAR Quartets for Amplitude Change Detection and Foundation-Model Evaluation
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- Knowledge-Guided Masked Autoencoder with Linear Spectral Mixing and Spectral-Angle-Aware Reconstruction
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural Activity
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
- Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- StainNet: Scaling Self-Supervised Foundation Models on Immunohistochemistry and Special Stains for Computational Pathology
- Quantifying Uncertainty in Machine Learning-Based Pervasive Systems: Application to Human Activity Recognition
- ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
- Masked Registration and Autoencoding of CT Images for Predictive Tibia Reconstruction
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
- Stanford Sleep Bench: Evaluating Polysomnography Pre-training Methods for Sleep Foundation Models
- Masked Generative Policy for Robotic Control
- Graph Deep Learning for Intracranial Aneurysm Blood Flow Simulation and Risk Assessment
- Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
- PointDico: Contrastive 3D Representation Learning Guided by Diffusion Models
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Towards Sustainable Universal Deepfake Detection with Frequency-Domain Masking
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- Improving action classification with brain-inspired deep networks
- How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline
- A Geometric Unification of Concept Learning with Concept Cones
- Radiance-Field Reinforced Pretraining: Scaling Localization Models with Unlabeled Wireless Signals
- Self-Supervised Learning on Molecular Graphs: A Systematic Investigation of Masking Design
- Masked Autoencoder Pretraining on Strong-Lensing Images for Joint Dark-Matter Model Classification and Super-Resolution
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- A Comparative Study on Synthetic Facial Data Generation Techniques for Face Recognition
- Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
- OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
- Self-Supervised Learning for Transparent Object Depth Completion Using Depth from Non-Transparent Objects
- Self-supervised prior learning improves structured illumination microscopy resolution
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
- The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- Boundary-Aware Test-Time Adaptation for Zero-Shot Medical Image Segmentation
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
- Vision and Causal Learning Based Channel Estimation for THz Communications
- Minuet: A Diffusion Autoencoder for Compact Semantic Compression of Multi-Band Galaxy Images
- Unique Lives, Shared World: Learning from Single-Life Videos
- On the Temporality for Sketch Representation Learning
- Diminishing Returns in Self-Supervised Learning
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
- Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features
- Vision Foundry: A System for Training Foundational Vision AI Models
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Unsupervised Structural Scene Decomposition via Foreground-Aware Slot Attention with Pseudo-Mask Guidance
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade
- PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- Lost in Distortion: Uncovering the Domain Gap Between Computer Vision and Brain Imaging -- A Study on Pretraining for Age Prediction
- nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
- First On-Orbit Demonstration of a Geospatial Foundation Model
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- Self-sufficient Independent Component Analysis via KL Minimizing Flows
- SelfAI: Building a Self-Training AI System with LLM Agents
- UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- PowerCLIP: Powerset Alignment for Contrastive Pre-Training
- Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- TS2Vec-Ensemble: An Enhanced Self-Supervised Framework for Time Series Forecasting
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
- Structure is Supervision: Multiview Masked Autoencoders for Radiology
- The Collapse of Patches
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Tiny-TSM: Efficiently Training a Lightweight SOTA Time Series Foundation Model
- Automated Histopathologic Assessment of Hirschsprung Disease Using a Multi-Stage Vision Transformer Framework
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation
- Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
- LungEvaty: A Scalable, Open-Source Transformer-based Deep Learning Model for Lung Cancer Risk Prediction in LDCT Screening
- Foundry: Distilling 3D Foundation Models for the Edge
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- Medal S: Spatio-Textual Prompt Model for Medical Segmentation
- DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
- Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
- Rethinking Vision Transformer Depth via Structural Reparameterization
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- stable-pretraining-v1: Foundation Model Research Made Simple
- SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- A cross-species neural foundation model for end-to-end speech decoding
- PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
- Quantum Masked Autoencoders for Vision Learning
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Dual-domain Adaptation Networks for Realistic Image Super-resolution
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Investigating self-supervised representations for audio-visual deepfake detection
- CroTad: A Contrastive Reinforcement Learning Framework for Online Trajectory Anomaly Detection
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- Generative Augmented Reality: Paradigms, Technologies, and Future Applications
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- Toward Artificial Palpation: Representation Learning of Touch on Soft Bodies
- Graph Neural Networks for Surgical Scene Segmentation
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- Crossmodal learning for Crop Canopy Trait Estimation
- SpectralTrain: A Universal Framework for Hyperspectral Image Classification
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- A Spatial Semantics and Continuity Perception Attention for Remote Sensing Water Body Change Detection
- UniSER: A Foundation Model for Unified Soft Effects Removal
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
- Segment Anything Across Shots: A Method and Benchmark
- Training-free Detection of AI-generated images via Cropping Robustness
- Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Quantum Machine Learning via Contrastive Training
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- Passive Dementia Screening via Facial Temporal Micro-Dynamics Analysis of In-the-Wild Talking-Head Video
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- An Evaluation of Representation Learning Methods in Particle Physics Foundation Models
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation
- Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Learning the relative composition of EEG signals using pairwise relative shift pretraining
- Binary Verification for Zero-Shot Vision
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- DermAI: Clinical dermatology acquisition through quality-driven image collection for AI classification in mobile
- Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training
- MIRNet: Integrating Constrained Graph-Based Reasoning with Pre-training for Diagnostic Medical Imaging
- SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- Leveraging unlabelled data for generalizable neural population decoding
- One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- Visual Bridge: Universal Visual Perception Representations Generating
- Deep Learning Analysis of Prenatal Ultrasound for Identification of Ventriculomegaly
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- SYNAPSE: Synergizing an Adapter and Finetuning for High-Fidelity EEG Synthesis from a CLIP-Aligned Encoder
- RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- LeCoT: revisiting network architecture for two-view correspondence pruning
- Learning from the Right Patches: A Two-Stage Wavelet-Driven Masked Autoencoder for Histopathology Representation Learning
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- CINEMAE: Leveraging Frozen Masked Autoencoders for Cross-Generator AI Image Detection
- CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purification for Inference-Time Attacks
- Frequency Matters: When Time Series Foundation Models Fail Under Spectral Shift
- Multimodal Diffusion Forcing for Forceful Manipulation
- An Active Learning Pipeline for Biomedical Image Instance Segmentation with Minimal Human Intervention
- Data Efficiency and Transfer Robustness in Biomedical Image Segmentation: A Study of Redundancy and Forgetting with Cellpose
- Cambrian-S: Towards Spatial Supersensing in Video
- Landslide Hazard Mapping with Geospatial Foundation Models: Geographical Generalizability, Data Scarcity, and Band Adaptability
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
- MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
- THD-BAR: Topology Hierarchical Derived Brain Autoregressive Modeling for EEG Generic Representations
- Unsupervised whole-heart function assessment
- Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
- An Augmentation Overlap Theory of Contrastive Learning
- A Foundation Model for Brain MRI with Dynamic Modality Integration
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- Dynamic Reflections: Probing Video Representations with Text Alignment
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- Differentiable Hierarchical Visual Tokenization
- Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework
- Diffusion Models are Robust Pretrainers
- PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
- Assessing the value of Geo-Foundational Models for Flood Inundation Mapping: Benchmarking models for Sentinel-1, Sentinel-2, and Planetscope for end-users
- Anatomically Constrained Transformers for Echocardiogram Analysis
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- Region-Aware Reconstruction Strategy for Pre-training fMRI Foundation Model
- Leveraging Generic Time Series Foundation Models for EEG Classification
- Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis
- SpecAware: A Spectral-Content Aware Foundation Model for Unifying Multi-Sensor Learning in Hyperspectral Remote Sensing Mapping
- A Step Toward World Models: A Survey on Robotic Manipulation
- Privacy-Aware Continual Self-Supervised Learning on Multi-Window Chest Computed Tomography for Domain-Shift Robustness
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras
- Constructing the Umwelt: Cognitive Planning through Belief-Intent Co-Evolution
- Robust Super-Capacity SRS Channel Inpainting via Diffusion Models
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- Active Learning with Task-Driven Representations for Messy Pools
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
- Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
- Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Natural Scene Text Editing Based on AI
- Bridging Global Context Interactions for High-Fidelity Image Completion
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- Amortized Moment Matching for Visual Generation
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Is Dimensionality a Barrier for Retrieval Models?
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Microscopy Images
- Zero-shot World Models Are Developmentally Efficient Learners
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- NeurIPT: Foundation Model for Neural Interfaces
- RL makes MLLMs see better than SFT
- Generative Modeling via Drifting
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
- Attentive multilayer fusion for vision transformers
- A multimodal whole-slide foundation model for pathology
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Eigenfunction Extraction for Ordered Representation Learning
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- A Unified Geometric Space Bridging AI Models and the Human Brain
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- PULSE: Privileged Knowledge Transfer from Electrodermal Activity to Low-Cost Sensors for Stress Monitoring
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoencoding
- SeeDNorm: Self-Rescaled Dynamic Normalization
- WaveMAE: Wavelet decomposition Masked Auto-Encoder for Remote Sensing
- DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- Mutual Information guided Visual Contrastive Learning
- Simplifying Knowledge Transfer in Pretrained Models
- Learning 3D Anisotropic Noise Distributions Improves Molecular Force Field Modeling
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
- VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Towards Objective Obstetric Ultrasound Assessment: Contrastive Representation Learning for Fetal Movement Detection
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Diffusion Autoencoders with Perceivers for Long, Irregular and Multimodal Astronomical Sequences
- Exploring Conditions for Diffusion models in Robotic Control
- Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
- Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization
- See the Text: From Tokenization to Visual Reading
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
- Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
- ShortcutBreaker: Low-Rank Noisy Bottleneck with Global Perturbation Attention for Multi-Class Unsupervised Anomaly Detection
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Post-Processing Methods for Improving Accuracy in MRI Inpainting
- OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales
- “UDE DIATOMS in the Wild 2024”: a new image dataset of freshwater diatoms for training deep learning models
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
- Confidence-Weighted Semi-Supervised Learning for Skin Lesion Segmentation Using Hybrid CNN-Transformer Networks
- NeuCo-Bench: A Novel Benchmark Framework for Neural Embeddings in Earth Observation
- Do Satellite Tasks Need Special Pretraining?
- 3D-GSRD: 3D Molecular Graph Auto-Encoder with Selective Re-mask Decoding
- Latent Diffusion Model without Variational Autoencoder
- Decorrelation Speeds Up Vision Transformers
- Comprehensive language-image pre-training for 3D medical image understanding
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
- Cognitive-Aligned Spatio-Temporal Large Language Models For Next Point-of-Interest Prediction
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- Towards Adversarial Robustness and Uncertainty Quantification in DINOv2-based Few-Shot Anomaly Detection
- Scaling Vision Transformers for Functional MRI with Flat Maps
- EEGChaT: A Transformer-Based Modular Channel Selector for SEEG Analysis
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- Transformer-based Scalable Beamforming Optimization via Deep Residual Learning
- Universal Image Restoration Pre-training via Masked Degradation Classification
- Reasoning in Space via Grounding in the World
- DP-TTA: Test-time Adaptation for Transient Electromagnetic Signal Denoising via Dictionary-driven Prior Regularization
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- CAMNet: Leveraging Cooperative Awareness Messages for Vehicle Trajectory Prediction
- AnyUp: Universal Feature Upsampling
- Pretraining in Actor-Critic Reinforcement Learning for Robot Locomotion
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- Diffusion Transformers with Representation Autoencoders
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis
- Inpainting the Neural Picture: Inferring Unrecorded Brain Area Dynamics from Multi-Animal Datasets
- Benchmarking foundation models for hyperspectral image classification: Application to cereal crop type mapping
- Exploring and Leveraging Class Vectors for Classifier Editing
- G2L:From Giga-Scale to Cancer-Specific Large-Scale Pathology Foundation Models via Knowledge Distillation
- DAWP: A framework for global observation forecasting via Data Assimilation and Weather Prediction in satellite observation space
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Redundancy as a Structural Information Principle for Learning and Generalization
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
- Visual Odometry with Transformers
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- Population-Coded Spiking Neural Networks for High-Dimensional Robotic Control
- Jigsaw3D: Disentangled 3D Style Transfer via Patch Shuffling and Masking
- PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
- SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
- Unsupervised Dynamic Feature Selection for Robust Latent Spaces in Vision Tasks
- Reducing Simulation Dependence in Neutrino Telescopes with Masked Point Transformers
- An uncertainty-aware framework for data-efficient multi-view animal pose estimation
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Neural Codecs as Biosignal Tokenizers
- SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense
- A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
- SatFusion: A Unified Framework for Enhancing Satellite IoT Images via Multi-Temporal and Multi-Source Data Fusion
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection
- DADO: A Depth-Attention framework for Object Discovery
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Adaptive Semantic Communication for UAV/UGV Cooperative Path Planning
- End-to-End Test-Time Training for Long Context
- Rapid computation of high-level visual surprise
- VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
- Latent Representation Learning in Heavy-Ion Collisions with MaskPoint Transformer
- Heptapod: Language Modeling on Visual Signals
- MSITrack: A Challenging Benchmark for Multispectral Single Object Tracking
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Dynamic Control Aware Semantic Communication Enabled Image Transmission for Lunar Landing
- AIM 2025 Challenge on Real-World RAW Image Denoising
- Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Conditional Representation Learning for Customized Tasks
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Glocal Information Bottleneck for Time Series Imputation
- Learning More with Less: A Generalizable, Self-Supervised Framework for Privacy-Preserving Capacity Estimation with EV Charging Data
- Mapping Rio de Janeiro's favelas: general-purpose vs. satellite-specific neural networks
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Cross-Modal Reconstruction Pretraining for Ramp Flow Prediction at Highway Interchanges
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- Creative synthesis of kinematic mechanisms
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training
- GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap Data
- The Impact of Scaling Training Data on Adversarial Robustness
- PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
- From Cheap Geometry to Expensive Physics: Elevating Neural Operators via Latent Shape Pretraining
- LieHMR: Autoregressive Human Mesh Recovery with SO(3) Diffusion
- Joint Embeddings Go Temporal
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- Agentic Services Computing
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Fast Feature Field (F3): A Predictive Representation of Events
- Personalized Vision via Visual In-Context Learning
- Training Agents Inside of Scalable World Models
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
- FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing
- Graph Your Own Prompt
- WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- Benchmarking DINOv3 for Multi-Task Stroke Analysis on Non-Contrast CT
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- Data-Efficient Training by Evolved Sampling
- Orochi: Versatile Biomedical Image Processor
- Category Discovery: An Open-World Perspective
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- Stochastic activations
- BrainPro: Towards Large-scale Brain State-aware EEG Representation Learning
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- On the Status of Foundation Models for SAR Imagery
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
- Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
- TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
- U-Mamba2-SSL for Semi-Supervised Tooth and Pulp Segmentation in CBCT
- Incomplete Data, Complete Dynamics: A Diffusion Approach
- Embodied AI: From LLMs to World Models
- Generalist Robot Manipulation beyond Action Labeled Data
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- Anatomically Constrained Transformers for Cardiac Amyloidosis Classification
- Shared Neural Space: Unified Precomputed Feature Encoding for Multi-Task and Cross Domain Vision
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- Are Foundation Models Ready for Industrial Defect Recognition? A Reality Check on Real-World Data
- Robust AI-ECG for Predicting Left Ventricular Systolic Dysfunction in Pediatric Congenital Heart Disease
- SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- Transporting Task Vectors across Different Architectures without Training
- SleepLM: Natural-Language Intelligence for Human Sleep
- OSF: On Pre-training and Scaling of Sleep Foundation Models
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- SSNet: Flexible and robust channel extrapolation for fluid antenna systems enabled by an self-supervised learning framework
- A2M2-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition
- Overview of PlantCLEF 2022: Image-based plant identification at global scale
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Multimodal Medical Image Classification via Synergistic Learning Pre-training
- SingLEM: Single-Channel Large EEG Model
- Overview of PlantCLEF 2022: Image-based plant identification at global scale
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- Mixture of Noise for Pre-Trained Model-Based Class-Incremental Learning
- Self-Supervised Learning of Graph Representations for Network Intrusion Detection
- ViTCAE: ViT-based Class-conditioned Autoencoder
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- The Missing Piece: A Case for Pre-Training in 3D Medical Object Detection
- Opportunities and Challenges in Applying AI to Evolutionary Morphology
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- Beyond Random Masking: A Dual-Stream Approach for Rotation-Invariant Point Cloud Masked Autoencoders
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- Designing Latent Safety Filters using Pre-Trained Vision Models
- CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
- Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
- AToken: A Unified Tokenizer for Vision
- Masked Feature Modeling Enhances Adaptive Segmentation
- Self Identity Mapping
- Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
- SAMIR, an efficient registration framework via robust feature learning from SAM
- CSMoE: An Efficient Remote Sensing Foundation Model with Soft Mixture-of-Experts
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- First Place Solution to the MLCAS 2025 GWFSS Challenge: The Devil is in the Detail and Minority
- Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
- Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
- Empowering Multi-Robot Cooperation via Sequential World Models
- Bridging Performance Gaps for ECG Foundation Models: A Post-Training Strategy
- Improving Anomalous Sound Detection with Attribute-aware Representation from Domain-adaptive Pre-training
- Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single Image
- FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Road Obstacle Video Segmentation
- Multi Anatomy X-Ray Foundation Model
- Image Tokenizer Needs Post-Training
- RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration
- A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications
- DRAG: Data Reconstruction Attack using Guided Diffusion
- Two-Stage Decoupling Framework for Variable-Length Glaucoma Prognosis
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder
- ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Lightweight Metadata-Aware Mixture-of-Experts Masked Autoencoder for Earth Observation
- Multimodal SAM-adapter for Semantic Segmentation
- SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- Improving Audio Event Recognition with Consistency Regularization
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge
- FLARE-SSM: Deep State Space Models with Influence-Balanced Loss for 72-Hour Solar Flare Prediction
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
- Improvement of Human-Object Interaction Action Recognition Using Scene Information and Multi-Task Learning Approach
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- Object-level Correlation for Few-Shot Segmentation
- HU-based Foreground Masking for 3D Medical Masked Image Modeling
- Self-Supervised Cross-Encoder for Neurodegenerative Disease Diagnosis
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
- Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis
- Curia: A Multi-Modal Foundation Model for Radiology
- SVGauge: Towards Human-Aligned Evaluation for SVG Generation
- MedSeqFT: Sequential Fine-tuning Foundation Models for 3D Medical Image Segmentation
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Towards Open World Detection: A Survey
- CPEP: Contrastive Pose-EMG Pre-training Enhances Gesture Generalization on EMG Signals
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- Parking Availability Prediction via Fusing Multi-Source Data with A Self-Supervised Learning Enhanced Spatio-Temporal Inverted Transformer
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- The Protocol Genome A Self Supervised Learning Framework from DICOM Headers
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
- Self-Validated Learning for Particle Separation: A Correctness-Based Self-Training Framework Without Human Labels
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Examination of PCA Utilisation for Multilabel Classifier of Multispectral Images
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- Multitask Battery Management with Flexible Pretraining
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
- Temporal Representation Learning for Real-Time Ultrasound Analysis
- Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
- CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition
- ER-LoRA: Effective-Rank Guided Adaptation for Weather-Generalized Depth Estimation
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Double-Constraint Diffusion Model with Nuclear Regularization for Ultra-low-dose PET Reconstruction
- CoMET: A Contrastive-Masked Brain Foundation Model for Universal EEG Representation
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- SatDINO: A Deep Dive into Self-Supervised Pretraining for Remote Sensing
- Representation Learning with Adaptive Superpixel Coding
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Generalizable Object Re-Identification via Visual In-Context Prompting
- GSTBench: A Benchmark Study on the Transferability of Graph Self-Supervised Learning
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- EEGDM: Learning EEG Representation with Latent Diffusion Model
- Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Occlusion Robustness of CLIP for Military Vehicle Classification
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- ECG-Soup: Harnessing Multi-Layer Synergy for ECG Foundation Models
- Patch Progression Masked Autoencoder with Fusion CNN Network for Classifying Evolution Between Two Pairs of 2D OCT Slices
- A Masked Representation Learning to Model Cardiac Functions Using Multiple Physiological Signals
- Deep Pre-trained Time Series Features for Tree Species Classification in the Dutch Forest Inventory
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- Training Transformers for Mesh-Based Simulations
- EventSSEG: Event-driven Self-Supervised Segmentation with Probabilistic Attention
- Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- CuMoLoS-MAE: A Masked Autoencoder for Remote Sensing Data Reconstruction
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
- Surya: Foundation Model for Heliophysics
- DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Impact of Clinical Image Quality on Efficient Foundation Model Finetuning
- FunduSegmenter: Leveraging the RETFound Foundation Model for Joint Optic Disc and Optic Cup Segmentation in Retinal Fundus Images
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
- APFL: Analytic Personalized Federated Learning via Dual-Stream Least Squares
- A Retrieval Augmented Spatio-Temporal Framework for Traffic Prediction
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection
- GeoMAE: Masking Representation Learning for Spatio-Temporal Graph Forecasting with Missing Values
- A Unified Contrastive-Generative Framework for Time Series Classification
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
- Diverse Teaching and Label Propagation for Generic Semi-Supervised Medical Image Segmentation
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking
- 3D Human Mesh Estimation from Single View RGBD
- Prompt-Guided Relational Reasoning for Social Behavior Understanding with Vision Foundation Models
- Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
- Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification
- CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
- Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- Large Model Driven Solar Activity AI Forecaster: A Scalable Dual Data-Model Framework
- LungSurg: A Generative AI System for Segmentation and Phase Classification in Thoracoscopic Lobectomy
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- impuTMAE: Multi-modal Transformer with Masked Pre-training for Missing Modalities Imputation in Cancer Survival Prediction
- DSConv: Dynamic Splitting Convolution for Pansharpening
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks
- Few-Shot Deployment of Pretrained MRI Transformers in Brain Imaging Tasks
- TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios
- Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- FedHiP: Heterogeneity-Invariant Personalized Federated Learning Through Closed-Form Solutions
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Benchmarking Foundation Models for Mitotic Figure Classification
- Unmasking Interstitial Lung Diseases: Leveraging Masked Autoencoders for Diagnosis
- VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
- A Foundation Model for DAS Signal Recognition and Visual Prompt Tuning of the Pre-trained Model for Downstream Tasks
- A Foundational Multi-Modal Model for Few-Shot Learning
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Elucidating the Role of Feature Normalization in IJEPA
- Whole-body Representation Learning For Competing Preclinical Disease Risk Assessment
- M3HL: Mutual Mask Mix with High-Low Level Feature Consistency for Semi-Supervised Medical Image Segmentation
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
- Model Recycling Framework for Multi-Source Data-Free Supervised Transfer Learning
- Context Guided Transformer Entropy Modeling for Video Compression
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- EvoVLMA: Evolutionary Vision-Language Model Adaptation
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- Multi-Operator Few-Shot Learning for Generalization Across PDE Families
- Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles
- IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
- STF: Shallow-Level Temporal Feedback to Enhance Spiking Transformers
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- Towards Robust Semantic Correspondence: A Benchmark and Insights
- Multimodal Referring Segmentation: A Survey
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- MINR: Implicit Neural Representations with Masked Image Modelling
- Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Temporally Consistent Unsupervised Segmentation for Mobile Robot Perception
- Low-Cost Test-Time Adaptation for Robust Video Editing
- Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
- Spatiodynamic inference using vision-based generative modelling
Discussions
- @bowang0911.bsky.social showed me this cool paper that I'd never read before, about Masked Autoencoder (MAE) for images. The idea: an image encoder encodes non-masked patches, followed by a decoder us [bsky, 3 points, 1 comments]
- Perhaps the biggest overhaul I've done to this so far, it replaces the feature extractor with one that follows this paper!
arxiv.org/abs/2111.06377 [bsky, 1 points, 0 comments]
- Original paper is this: arxiv.org/abs/2111.06377. They propose a couple of intuitive similarities between vision and language and explore what has been done in language with so much success but applie [bsky, 1 points, 1 comments]
- or masking techniques where you try to inpaint missing parts of a picture. For example MAE (Masked AutoEncoders https://arxiv.org/abs/2111.06377) [bsky, 0 points, 1 comments]
Related