Exploring Simple Siamese Representation Learning
2021/06/01 by Xinlei Chen, Kaiming He · 299 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications
paper · doi:10.1109/cvpr46437.2021.01549
openalex publication_date 2021/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Siamese networks have become a common structure in various recent models for unsupervised visual representation learning. These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing. We provide a hypothesis on the implication of stop-gradient, and further show proof-of-concept experiments verifying it. Our "SimSiam" method achieves competitive results on ImageNet and downstream tasks. We hope this simple baseline will motivate people to rethink the roles of Siamese architectures for unsupervised representation learning. Code is made available.1
Citations
Cited by
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Cross-Receiver Radio Frequency Fingerprint Identification Based on Contrastive Learning and Subdomain Adaptation
- An Invitation to Deep Reinforcement Learning
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Pre-training Molecular Graph Representation with 3D Geometry
- GenURL: A General Framework for Unsupervised Representation Learning
- Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
- RL makes MLLMs see better than SFT
- Generative Modeling via Drifting
- Eigenfunction Extraction for Ordered Representation Learning
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame Projections
- Mutual Information guided Visual Contrastive Learning
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- Scaling Non-Parametric Sampling with Representation
- Interpretable Multimodal Zero-Shot ECG Diagnosis via Structured Clinical Knowledge Alignment
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- Resounding Acoustic Fields with Reciprocity
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Ego-Vision World Model for Humanoid Contact Planning
- Contrastive Noise-Guided Invertible Network for Image Steganography
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- SpurBreast: A Curated Dataset for Investigating Spurious Correlations in Real-world Breast MRI Classification
- Median2Median: Zero-shot Suppression of Structured Noise in Images
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Contrastive Self-Supervised Learning at the Edge: An Energy Perspective
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Conditional Representation Learning for Customized Tasks
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Glocal Information Bottleneck for Time Series Imputation
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
- Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
- Diverse Text-to-Image Generation via Contrastive Noise Optimization
- Self-Supervised Representation Learning as Mutual Information Maximization
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Targeted Supervised Contrastive Learning for Long-Tailed Recognition
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- A Family of Kernelized Matrix Costs for Multiple-Output Mixture Neural Networks
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- An Investigation into the Performance of Non-Contrastive Self-Supervised Learning Methods for Network Intrusion Detection
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse
- Generation Properties of Stochastic Interpolation under Finite Training Set
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- Bootstrapping User and Item Representations for One-Class Collaborative Filtering
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- Enhancing hyperspectral image prediction with contrastive learning in low-label regimes
- Temporal Straightening for Latent Planning
- Few-shot Human Action Anomaly Detection via a Unified Contrastive Learning Framework
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- A Novel Task-Driven Diffusion-Based Policy with Affordance Learning for Generalizable Manipulation of Articulated Objects
- Self Identity Mapping
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Text-Based Person Search with Limited Data
- Improving Audio Event Recognition with Consistency Regularization
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Fed-REACT: Federated Representation Learning for Heterogeneous and Evolving Data
- Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- Rethinking Supervised Pre-training for Better Downstream Transferring
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- The Protocol Genome A Self Supervised Learning Framework from DICOM Headers
- LLMCDSR: Enhancing Cross-Domain Sequential Recommendation with Large Language Models
- A Metaverse: Taxonomy, Components, Applications, and Open Challenges
- Examination of PCA Utilisation for Multilabel Classifier of Multispectral Images
- Domain Adaptation via Feature Refinement
- Self-Supervised Learning for Gastritis Detection with Gastric X-ray Images
- Decomposing and Revising What Language Models Generate
- Bag of Tricks and A Strong baseline for Image Copy Detection
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- Learning from Silence and Noise for Visual Sound Source Localization
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- DNP-Guided Contrastive Reconstruction with a Reverse Distillation Transformer for Medical Anomaly Detection
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- A Generalized Learning Framework for Self-Supervised Contrastive Learning
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Learning Representations for Pixel-based Control: What Matters and Why?
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
- DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
- Bootstrap Deep Spectral Clustering with Optimal Transport
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Bypassing Skip-Gram Negative Sampling: Dimension Regularization as a More Efficient Alternative for Graph Embeddings
- PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
- Learning Deep Representation with Energy-Based Self-Expressiveness for Subspace Clustering
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
- Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
- Consistency Regularization for Deep Face Anti-Spoofing
- Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements
- Ensemble Foreground Management for Unsupervised Object Discovery
- Style-Aware Blending and Prototype-Based Cross-Contrast Consistency for Semi-Supervised Medical Image Segmentation
- SpecBPP: A Self-Supervised Learning Approach for Hyperspectral Representation and Soil Organic Carbon Estimation
- Extreme Cardiac MRI Analysis under Respiratory Motion: Results of the CMRxMotion Challenge
- MC-SSL0.0: Towards Multi-Concept Self-Supervised Learning
- Leveraging Data Augmentation and Siamese Learning for Predictive Process Monitoring
- Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images
- Improving Contrastive Learning by Visualizing Feature Transformation
- MixSiam: A Mixture-based Approach to Self-supervised Representation Learning
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- C3RL: Rethinking the Combination of Channel-independence and Channel-mixing from Representation Learning
- Exploring Active Learning for Semiconductor Defect Segmentation
- Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
- Dataset Ownership Verification for Pre-trained Masked Models
- On the Similarities of Embeddings in Contrastive Learning
- Emerging Properties in Self-Supervised Vision Transformers
- Cluster Contrast for Unsupervised Visual Representation Learning
- From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
- Self-supervised predictive learning accounts for cortical layer-specificity
- The Impact of Spatiotemporal Augmentations on Self-Supervised Audiovisual Representation Learning
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- Diffuse and Disperse: Image Generation with Representation Regularization
- Clustering-Guided Multi-Layer Contrastive Representation Learning for Citrus Disease Classification
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
- Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
- PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies
- Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation
- Self-Supervised Learning at the Edge: The Cost of Labeling
- CL-Polyp: A Contrastive Learning-Enhanced Network for Accurate Polyp Segmentation
- Omni-Fusion of Spatial and Spectral for Hyperspectral Image Segmentation
- Divergence-Based Similarity Function for Multi-View Contrastive Learning
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- GIST: Cross-Domain Click-Through Rate Prediction via Guided Content-Behavior Distillation
- Clustering via Self-Supervised Diffusion
- Deconfounding Causal Inference through Two-Branch Framework with Early-Forking for Sensor-Based Cross-Domain Activity Recognition
- Large-Scale Hyperspectral Image Clustering Using Contrastive Learning
- IGDNet: Zero-Shot Robust Underexposed Image Enhancement via Illumination-Guided and Denoising
- Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
- MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping
- Compressive Visual Representations
- Unsupervised Representation Learning for Binary Networks by Joint Classifier Learning
- PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View
- Emergent musical properties of a transformer under contrastive self-supervised learning
- A Theoretical Formulation on the Use of Multiple Positive Views in Contrastive Learning
- Self-Supervised Contrastive Learning for Multi-Label Images
- Lightweight Physics-Aware Zero-Shot Ultrasound Plane-Wave Denoising
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical Vision
- Multiple Object Stitching for Unsupervised Representation Learning
- Resampling Augmentation for Time Series Contrastive Learning: Application to Remote Sensing
- Object-aware Sound Source Localization via Audio-Visual Scene Understanding
- Leveraging neural network interatomic potentials for a foundation model of chemistry
- DIP: Unsupervised Dense In-Context Post-training of Visual Representations
- Weakly-supervised Generative Adversarial Networks for medical image classification
- DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
- Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing
- SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
- A Survey of State Representation Learning for Deep Reinforcement Learning
- Bridging Brain with Foundation Models through Self-Supervised Learning
- Cluster Analysis with Deep Embeddings and Contrastive Learning
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
- Contrastive Self-Supervised Learning As Neural Manifold Packing
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- Enhancing Large Language Models for Mobility Analytics with Semantic Location Tokenization
- Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports
- Task-Driven Discrete Representation Learning
- Visual Pre-Training on Unlabeled Images using Reinforcement Learning
- A Prototype-Oriented Framework for Unsupervised Domain Adaptation
- Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks
- Rethinking Graph Contrastive Learning through Relative Similarity Preservation
- Exploring Visual Prompting: Robustness Inheritance and Beyond
- When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
- QualitEye: Public and Privacy-preserving Gaze Data Quality Verification
- Rethinking Semi-supervised Segmentation Beyond Accuracy: Reliability and Robustness
- Self-supervised One-Stage Learning for RF-based Multi-Person Pose Estimation
- FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
- OWT: A Foundational Organ-Wise Tokenization Framework for Medical Imaging
- Language-Image Alignment with Fixed Text Encoders
- Unleashing the Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-Identification
- How PARTs assemble into wholes: Learning the relative composition of images
- ConMamba: Contrastive Vision Mamba for Plant Disease Detection
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Aligned Contrastive Loss for Long-Tailed Recognition
- anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- HMAE: Self-Supervised Few-Shot Learning for Quantum Spin Systems
- SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Editing
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
- Online Unsupervised Learning of Visual Representations and Categories
- Diversify and Conquer: Open-set Disagreement for Robust Semi-supervised Learning with Outliers
- A Mathematical Perspective On Contrastive Learning
- A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
- A Tutorial-cum-Survey on Self-Supervised Learning for Wi-Fi Sensing: Trends, Challenges, and Outlook
- Self-supervised feature learning for cardiac Cine MR image reconstruction
- FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
- Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
- D2LV: A Data-Driven and Local-Verification Approach for Image Copy Detection
- An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
- Information-Theoretic Complementary Prompts for Improved Continual Text Classification
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- Self-Supervised Evolution Operator Learning for High-Dimensional Dynamical Systems
- Generative AI and foundation models in medical image
- Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- Simple Contrastive Representation Adversarial Learning for NLP Tasks
- Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning
- InfoGCL: Information-Aware Graph Contrastive Learning
- CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
- Utilizing Strategic Pre-training to Reduce Overfitting: Baguan -- A Pre-trained Weather Forecasting Model
- Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
- Collaborative Unlabeled Data Optimization
- A Few Large Shifts: Layer-Inconsistency Based Minimal Overhead Adversarial Example Detection
- Self-Supervised Learning for Image Segmentation: A Comprehensive Survey
- scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell Data
- AdaDim: Dimensionality Adaptation for SSL Representational Dynamics
- Spectral-Spatial Self-Supervised Learning for Few-Shot Hyperspectral Image Classification
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
- Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios
- SECRET: Semi-supervised Clinical Trial Document Similarity Search
- Logographic Character Visual Pretraining via Semantic-based Contrastive Learning
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
- AROpt: An Optimization Method for Autoregressive Time Series Forecasting
- Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
- Enhancing User Sequence Modeling through Barlow Twins-based Self-Supervised Learning
- FreCT: Frequency-augmented Convolutional Transformer for Robust Time Series Anomaly Detection
- AdaJEPA: An Adaptive Latent World Model
- Self-Supervised Learning by Estimating Twin Class Distributions
- RIO: Rotation-equivariance supervised learning of robust inertial odometry
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- When to Align, When to Predict: A Phase Diagram for Multimodal Learning
- Color Variants Identification in Fashion e-commerce via Contrastive Self-Supervised Representation Learning
- Next Embedding Prediction Makes World Models Stronger
- StreamSplit: Continuous Audio Representation Learning via Uncertainty-Guided Adaptive Splitting
- When Does LeJEPA Learn a World Model?
- Exploring internal representation of self-supervised networks: few-shot learning abilities and comparison with human semantics and recognition of objects
- Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
- MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis
- Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
- ASMa: Asymmetric Spatio-temporal Masking for Skeleton Action Representation Learning
- Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
- A Genealogy of Foundation Models in Remote Sensing
- NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
- When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
- Robust Semantic Segmentation with Superpixel-Mix
- Representation Learning via Non-Contrastive Mutual Information
- Variational Self-Supervised Learning
- Anomaly Detection on Attributed Networks via Contrastive Self-Supervised Learning
- Representation Synthesis by Probabilistic Many-Valued Logic Operation in Self-Supervised Learning
- Multimodal Perception for Goal-oriented Navigation: A Survey
- Boosting Generative Image Modeling via Joint Image-Feature Synthesis
- SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-training
- Can We Ignore Labels In Out of Distribution Detection?
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- Decomposition-based multi-scale transformer framework for time series anomaly detection
- DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
- Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-Supervised Speaker Verification
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
- Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data for User Perception
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
- Evolved Hierarchical Masking for Self-Supervised Learning
- JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
- Impact of Language Guidance: A Reproducibility Study
- Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
- Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
- SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
Related