Exploring Simple Siamese Representation Learning
2020/11/20 by Xinlei Chen, Kaiming He, Chen, Xinlei +1 · 368 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2011.10566
Technical report, 10 pages
arxiv created 2020/11/20 · openalex publication_date 2020/11/20 · arxiv updated 2020/11/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Siamese networks have become a common structure in various recent models for unsupervised visual representation learning. These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing. We provide a hypothesis on the implication of stop-gradient, and further show proof-of-concept experiments verifying it. Our "SimSiam" method achieves competitive results on ImageNet and downstream tasks. We hope this simple baseline will motivate people to rethink the roles of Siamese architectures for unsupervised representation learning. Code will be made available.
Citations
Cited by
- Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- DA-SSL: self-supervised domain adaptor to leverage foundational models in turbt histopathology slides
- PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- Supervised Contrastive Frame Aggregation for Video Representation Learning
- SSA3D: Text-Conditioned Assisted Self-Supervised Framework for Automatic Dental Abutment Design
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- Self-Supervised Learning with Gaussian Processes
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception
- Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation
- Silhouette-based Gait Foundation Model
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
- You Can Trust Your Clustering Model: A Parameter-free Self-Boosting Plug-in for Deep Clustering
- Pre-train to Gain: Robust Learning Without Clean Labels
- Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing
- In Search of Goodness: Large Scale Benchmarking of Goodness Functions for the Forward-Forward Algorithm
- Esim: EVM Bytecode Similarity Detection Based on Stable-Semantic Graph
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- An Evaluation of Representation Learning Methods in Particle Physics Foundation Models
- Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
- PCA++: How Uniformity Induces Robustness to Background Noise in Contrastive Learning
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Φeat: Physically Grounded Material Feature Representation
- Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures
- Planning in Branch-and-Bound: Model-Based Reinforcement Learning for Exact Combinatorial Optimization
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
- Image-Intrinsic Priors for Integrated Circuit Defect Detection and Novel Class Discovery via Self-Supervised Learning
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- ANCHOR: Integrating Adversarial Training with Hard-mined Supervised Contrastive Learning for Robust Representation Learning
- Barlow Twins for Sequential Recommendation
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Contrastive Feature Loss for Image Prediction
- Imbalance-Aware Self-Supervised Learning for 3D Radiomic Representations
- RL makes MLLMs see better than SFT
- Generative Modeling via Drifting
- Eigenfunction Extraction for Ordered Representation Learning
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame Projections
- Mutual Information guided Visual Contrastive Learning
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- Scaling Non-Parametric Sampling with Representation
- Interpretable Multimodal Zero-Shot ECG Diagnosis via Structured Clinical Knowledge Alignment
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- Resounding Acoustic Fields with Reciprocity
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Contrastive Adaptive Propagation Graph Neural Networks for Efficient Graph Learning
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
- On Feature Decorrelation in Self-Supervised Learning
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Ego-Vision World Model for Humanoid Contact Planning
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- SpurBreast: A Curated Dataset for Investigating Spurious Correlations in Real-world Breast MRI Classification
- Median2Median: Zero-shot Suppression of Structured Noise in Images
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Contrastive Self-Supervised Learning at the Edge: An Energy Perspective
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Conditional Representation Learning for Customized Tasks
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Glocal Information Bottleneck for Time Series Imputation
- Diverse Text-to-Image Generation via Contrastive Noise Optimization
- Self-Supervised Representation Learning as Mutual Information Maximization
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- Tasting the cake: evaluating self-supervised generalization on out-of-distribution multimodal MRI data
- Enhancing hyperspectral image prediction with contrastive learning in low-label regimes
- A Family of Kernelized Matrix Costs for Multiple-Output Mixture Neural Networks
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- An Investigation into the Performance of Non-Contrastive Self-Supervised Learning Methods for Network Intrusion Detection
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Generation Properties of Stochastic Interpolation under Finite Training Set
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- Hard Negative Mixing for Contrastive Learning
- Temporal Straightening for Latent Planning
- Understanding Self-supervised Learning with Dual Deep Networks
- Self-Supervised Learning with Kernel Dependence Maximization
- BEiT: BERT Pre-Training of Image Transformers
- Few-shot Human Action Anomaly Detection via a Unified Contrastive Learning Framework
- DyFormer: A Scalable Dynamic Graph Transformer with Provable Benefits on Generalization Ability
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- A Novel Task-Driven Diffusion-Based Policy with Affordance Learning for Generalizable Manipulation of Articulated Objects
- Unsupervised Object-Level Representation Learning from Scene Images
- Self Identity Mapping
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Time-Series Representation Learning via Temporal and Contextual Contrasting
- Improving Audio Event Recognition with Consistency Regularization
- Contrastive Learning with Stronger Augmentations
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Fed-REACT: Federated Representation Learning for Heterogeneous and Evolving Data
- Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- The Protocol Genome A Self Supervised Learning Framework from DICOM Headers
- Examination of PCA Utilisation for Multilabel Classifier of Multispectral Images
- Domain Adaptation via Feature Refinement
- Contrastive Model Inversion for Data-Free Knowledge Distillation
- Bi-Granularity Contrastive Learning for Post-Training in Few-Shot Scene
- Decomposing and Revising What Language Models Generate
- Self-Guided Contrastive Learning for BERT Sentence Representations
- Learning from Silence and Noise for Visual Sound Source Localization
- Whitening for Self-Supervised Representation Learning
- DNP-Guided Contrastive Reconstruction with a Reverse Distillation Transformer for Medical Anomaly Detection
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- A Generalized Learning Framework for Self-Supervised Contrastive Learning
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
- SSFL: Tackling Label Deficiency in Federated Learning via Personalized Self-Supervision
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- FedCV: A Federated Learning Framework for Diverse Computer Vision Tasks
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
- DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
- Bootstrap Deep Spectral Clustering with Optimal Transport
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Bypassing Skip-Gram Negative Sampling: Dimension Regularization as a More Efficient Alternative for Graph Embeddings
- PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
- Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
- Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements
- Ensemble Foreground Management for Unsupervised Object Discovery
- Unsupervised Visual Representation Learning by Online Constrained K-Means
- Style-Aware Blending and Prototype-Based Cross-Contrast Consistency for Semi-Supervised Medical Image Segmentation
- SpecBPP: A Self-Supervised Learning Approach for Hyperspectral Representation and Soil Organic Carbon Estimation
- SVMax: A Feature Embedding Regularizer
- Extreme Cardiac MRI Analysis under Respiratory Motion: Results of the CMRxMotion Challenge
- Leveraging Data Augmentation and Siamese Learning for Predictive Process Monitoring
- Self-Supervised Neural Architecture Search for Imbalanced Datasets
- TENET: A Time-reversal Enhancement Network for Noise-robust ASR
- Dense Semantic Contrast for Self-Supervised Visual Representation Learning
- Jointly Learnable Data Augmentations for Self-Supervised GNNs
- DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
- ReSSL: Relational Self-Supervised Learning with Weak Augmentation
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- C3RL: Rethinking the Combination of Channel-independence and Channel-mixing from Representation Learning
- Exploring Active Learning for Semiconductor Defect Segmentation
- Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
- Dataset Ownership Verification for Pre-trained Masked Models
- On the Similarities of Embeddings in Contrastive Learning
- Emerging Properties in Self-Supervised Vision Transformers
- Towards Universal Dense Retrieval for Open-domain Question Answering
- Cluster Contrast for Unsupervised Visual Representation Learning
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
- CoDiM: Learning with Noisy Labels via Contrastive Semi-Supervised Learning
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- Pre-Trained Models: Past, Present and Future
- Diffuse and Disperse: Image Generation with Representation Regularization
- AlphaMatch: Improving Consistency for Semi-supervised Learning with Alpha-divergence
- Clustering-Guided Multi-Layer Contrastive Representation Learning for Citrus Disease Classification
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- Robust Contrastive Learning Using Negative Samples with Diminished Semantics
- Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
- Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
- PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies
- Self-Supervised Learning at the Edge: The Cost of Labeling
- CL-Polyp: A Contrastive Learning-Enhanced Network for Accurate Polyp Segmentation
- Omni-Fusion of Spatial and Spectral for Hyperspectral Image Segmentation
- Divergence-Based Similarity Function for Multi-View Contrastive Learning
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- GIST: Cross-Domain Click-Through Rate Prediction via Guided Content-Behavior Distillation
- Clustering via Self-Supervised Diffusion
- Improving Context-Based Meta-Reinforcement Learning with Self-Supervised Trajectory Contrastive Learning
- Deconfounding Causal Inference through Two-Branch Framework with Early-Forking for Sensor-Based Cross-Domain Activity Recognition
- A Note on Connecting Barlow Twins with Negative-Sample-Free Contrastive Learning
- Self-supervised predictive learning accounts for cortical layer-specificity
- IGDNet: Zero-Shot Robust Underexposed Image Enhancement via Illumination-Guided and Denoising
- Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
- MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View
- Emergent musical properties of a transformer under contrastive self-supervised learning
- Self-Supervised Contrastive Learning for Multi-Label Images
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Fighting Fire with Fire: Contrastive Debiasing without Bias-free Data via Generative Bias-transformation
- Lightweight Physics-Aware Zero-Shot Ultrasound Plane-Wave Denoising
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical Vision
- Multiple Object Stitching for Unsupervised Representation Learning
- Resampling Augmentation for Time Series Contrastive Learning: Application to Remote Sensing
- Object-aware Sound Source Localization via Audio-Visual Scene Understanding
- Leveraging neural network interatomic potentials for a foundation model of chemistry
- DIP: Unsupervised Dense In-Context Post-training of Visual Representations
- DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
- Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing
- SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
- A Survey of State Representation Learning for Deep Reinforcement Learning
- Bridging Brain with Foundation Models through Self-Supervised Learning
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
- Contrastive Self-Supervised Learning As Neural Manifold Packing
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
- SimTriplet: Simple Triplet Representation Learning with a Single GPU
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Enhancing Large Language Models for Mobility Analytics with Semantic Location Tokenization
- Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports
- Task-Driven Discrete Representation Learning
- Visual Pre-Training on Unlabeled Images using Reinforcement Learning
- Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks
- Rethinking Graph Contrastive Learning through Relative Similarity Preservation
- Exploring Visual Prompting: Robustness Inheritance and Beyond
- When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
- Homography augumented momentum constrastive learning for SAR image retrieval
- QualitEye: Public and Privacy-preserving Gaze Data Quality Verification
- Rethinking Semi-supervised Segmentation Beyond Accuracy: Reliability and Robustness
- Self-supervised One-Stage Learning for RF-based Multi-Person Pose Estimation
- FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
- OWT: A Foundational Organ-Wise Tokenization Framework for Medical Imaging
- Language-Image Alignment with Fixed Text Encoders
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
- How PARTs assemble into wholes: Learning the relative composition of images
- ConMamba: Contrastive Vision Mamba for Plant Disease Detection
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Aligned Contrastive Loss for Long-Tailed Recognition
- anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- HMAE: Self-Supervised Few-Shot Learning for Quantum Spin Systems
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Diversify and Conquer: Open-set Disagreement for Robust Semi-supervised Learning with Outliers
- A Mathematical Perspective On Contrastive Learning
- A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
- A Tutorial-cum-Survey on Self-Supervised Learning for Wi-Fi Sensing: Trends, Challenges, and Outlook
- Self-supervised feature learning for cardiac Cine MR image reconstruction
- FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
- Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
- Voxel-level Siamese Representation Learning for Abdominal Multi-Organ Segmentation
- An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
- Long-Short Temporal Contrastive Learning of Video Transformers
- Learning Rich Nearest Neighbor Representations from Self-supervised Ensembles
- Information-Theoretic Complementary Prompts for Improved Continual Text Classification
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
- Self-Supervised Evolution Operator Learning for High-Dimensional Dynamical Systems
- Generative AI and foundation models in medical image
- Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
- Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning
- An Empirical Study of Graph Contrastive Learning
- Visualizing MuZero Models
- CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
- An Empirical Study and Analysis on Open-Set Semi-Supervised Learning
- Learning From Long-Tailed Data With Noisy Labels
- Utilizing Strategic Pre-training to Reduce Overfitting: Baguan -- A Pre-trained Weather Forecasting Model
- Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
- Collaborative Unlabeled Data Optimization
- Momentum2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
- A Few Large Shifts: Layer-Inconsistency Based Minimal Overhead Adversarial Example Detection
- Self-Supervised Learning for Image Segmentation: A Comprehensive Survey
- scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell Data
- AdaDim: Dimensionality Adaptation for SSL Representational Dynamics
- Unsupervised Document Embedding via Contrastive Augmentation
- Spectral-Spatial Self-Supervised Learning for Few-Shot Hyperspectral Image Classification
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
- Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement
- Self-supervised Remote Sensing Images Change Detection at Pixel-level
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios
- SECRET: Semi-supervised Clinical Trial Document Similarity Search
- Logographic Character Visual Pretraining via Semantic-based Contrastive Learning
- Voxel-wise Cross-Volume Representation Learning for 3D Neuron Reconstruction
- Scale-Equivariant Imaging: Self-Supervised Learning for Image Super-Resolution and Deblurring
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- AROpt: An Optimization Method for Autoregressive Time Series Forecasting
- Enhancing User Sequence Modeling through Barlow Twins-based Self-Supervised Learning
- FreCT: Frequency-augmented Convolutional Transformer for Robust Time Series Anomaly Detection
- AdaJEPA: An Adaptive Latent World Model
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- When to Align, When to Predict: A Phase Diagram for Multimodal Learning
- Next Embedding Prediction Makes World Models Stronger
- StreamSplit: Continuous Audio Representation Learning via Uncertainty-Guided Adaptive Splitting
- When Does LeJEPA Learn a World Model?
- Exploring internal representation of self-supervised networks: few-shot learning abilities and comparison with human semantics and recognition of objects
- Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
- MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis
- Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
- ASMa: Asymmetric Spatio-temporal Masking for Skeleton Action Representation Learning
- Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
- A Genealogy of Foundation Models in Remote Sensing
- NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
- When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
- Representation Learning via Non-Contrastive Mutual Information
- Variational Self-Supervised Learning
- Anomaly Detection on Attributed Networks via Contrastive Self-Supervised Learning
- Self-Supervised Multisensor Change Detection
- Multimodal Perception for Goal-oriented Navigation: A Survey
- Boosting Generative Image Modeling via Joint Image-Feature Synthesis
- SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-training
- Can We Ignore Labels In Out of Distribution Detection?
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- Decomposition-based multi-scale transformer framework for time series anomaly detection
- DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
- IB-DRR: Incremental Learning with Information-Back Discrete Representation Replay
- Self-supervised Graph Neural Networks without explicit negative sampling
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
- Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data for User Perception
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
- Evolved Hierarchical Masking for Self-Supervised Learning
- JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
- Impact of Language Guidance: A Reproducibility Study
- Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
- Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
- SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
- Towards Generalizing Temporal Action Segmentation to Unseen Views
Related