Momentum Contrast for Unsupervised Visual Representation Learning
2020/06/01 by Kaiming He, Haoqi Fan, Yuxin Wu +2 · 484 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications
paper · doi:10.1109/cvpr42600.2020.00975
openalex publication_date 2020/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.
Citations
Cited by
- Self-Supervised Graph Co-Training for Session-based Recommendation
- ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- Pri3D: Can 3D Priors Help 2D Representation Learning?
- X-model: Improving Data Efficiency in Deep Learning with A Minimax Model
- Deep Long-Tailed Learning: A Survey
- Do Different Tracking Tasks Require Different Appearance Models?
- Unsupervised Natural Language Inference via Decoupled Multimodal Contrastive Learning
- PGL: Prior-Guided Local Self-supervised Learning for 3D Medical Image Segmentation
- Self-Supervised Multi-View Learning via Auto-Encoding 3D Transformations
- Mine Your Own vieW: Self-Supervised Learning Through Across-Sample\n Prediction
- Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift
- Self-supervised Heterogeneous Graph Neural Network with Co-contrastive Learning
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identification
- CLCC: Contrastive Learning for Color Constancy
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Deep semi-supervised learning for medical image segmentation: A review
- 3D Human Pose, Shape and Texture from Low-Resolution Images and Videos
- Characterizing signal propagation to close the performance gap in\n unnormalized ResNets
- SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
- Deep Learning with Label Differential Privacy
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Uncovering the structure of clinical EEG signals with self-supervised\n learning
- Pairwise Supervised Contrastive Learning of Sentence Representations
- Self-supervised Contrastive Video-Speech Representation Learning for Ultrasound
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- High-Performance Large-Scale Image Recognition Without Normalization
- Sparsity-Probe: Analysis tool for Deep Learning Models
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- Vector-quantized Image Modeling with Improved VQGAN
- Video Understanding as Machine Translation
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Deep Learning for Person Re-Identification: A Survey and Outlook
- Contrastive Feature Loss for Image Prediction
- Contrastive Semi-Supervised Learning for 2D Medical Image Segmentation
- Dual Contrastive Learning for Unsupervised Image-to-Image Translation
- Pre-training Molecular Graph Representation with 3D Geometry
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
- GenURL: A General Framework for Unsupervised Representation Learning
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- Trustworthy AI: From Principles to Practices
- Fundamental Limits and Tradeoffs in Invariant Representation Learning
- Automatic Shortcut Removal for Self-Supervised Representation Learning
- Imbalance-Aware Self-Supervised Learning for 3D Radiomic Representations
- CLDA: Contrastive Learning for Semi-Supervised Domain Adaptation
- Learning to Prompt for Vision-Language Models
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Exploring Data-Efficient 3D Scene Understanding with Contrastive Scene Contexts
- Cycle-Contrast for Self-Supervised Video Representation Learning
- RL makes MLLMs see better than SFT
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- Eigenfunction Extraction for Ordered Representation Learning
- Graph Neural Networks: Methods, Applications, and Opportunities
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Image Quality Assessment Using Contrastive Learning
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Domain Adaptive Semantic Segmentation with Self-Supervised Depth Estimation
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- Mutual Information guided Visual Contrastive Learning
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
- Deep Graph Contrastive Representation Learning
- Unsupervised Part Discovery from Contrastive Reconstruction
- Contrastive Learning for Many-to-many Multilingual Neural Machine Translation
- Resounding Acoustic Fields with Reciprocity
- Exploring Conditions for Diffusion models in Robotic Control
- MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
- Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Towards the Generalization of Contrastive Self-Supervised Learning
- CIL: Contrastive Instance Learning Framework for Distantly Supervised Relation Extraction
- ε-Seg: Sparsely Supervised Semantic Segmentation of Microscopy Data
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Self-Supervised Longitudinal Neighbourhood Embedding
- Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
- Token-Level Inference-Time Alignment for Vision-Language Models
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- Confidence-Weighted Semi-Supervised Learning for Skin Lesion Segmentation Using Hybrid CNN-Transformer Networks
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Geospatial Machine Learning Libraries
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Latent Diffusion Model without Variational Autoencoder
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Morphology-Aware Prognostic model for Five-Year Survival Prediction in Colorectal Cancer from H&E Whole Slide Images
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- ViTacGen: Robotic Pushing with Vision-to-Touch Generation
- Progressive Cluster Purification for Unsupervised Feature Learning
- Rethinking Graph Domain Adaptation: A Spectral Contrastive Perspective
- Universal Image Restoration Pre-training via Masked Degradation Classification
- Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation
- On Feature Decorrelation in Self-Supervised Learning
- Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
- FedHUG: Federated Heterogeneous Unsupervised Generalization for Remote Physiological Measurements
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- Diffusion Transformers with Representation Autoencoders
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Are Labels Necessary for Neural Architecture Search?
- Feature Stylization and Domain-aware Contrastive Learning for Domain Generalization
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Contrastive Dimension Reduction: A Systematic Review
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Scaling Language-Centric Omnimodal Representation Learning
- Contrastive Noise-Guided Invertible Network for Image Steganography
- Rectifying the Shortcut Learning of Background for Few-Shot Learning
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- Redundancy as a Structural Information Principle for Learning and Generalization
- A Joint Learning Approach to Hardware Caching and Prefetching
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- Decoupling Representation Learning from Reinforcement Learning
- Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
- Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss
- Astronomia ex machina: a history, primer, and outlook on neural networks in astronomy
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- 3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Vision Language Models: A Survey of 26K Papers
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- Conditional Alignment and Uniformity for Contrastive Learning with\n Continuous Proxy Labels
- Contrastive Representations for Label Noise Require Fine-Tuning
- A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Batch Curation for Unsupervised Contrastive Representation Learning
- Contrastive Self-Supervised Learning at the Edge: An Energy Perspective
- Label-Efficient Multi-Task Segmentation using Contrastive Learning
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- DADO: A Depth-Attention framework for Object Discovery
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
- SSL-SE-EEG: A Framework for Robust Learning from Unlabeled EEG Data with Self-Supervised Learning and Squeeze-Excitation Networks
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- Conditional Representation Learning for Customized Tasks
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Contrastive Representation Regularization for Vision-Language-Action Models
- Cross-Batch Negative Sampling for Training Two-Tower Recommenders
- Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph Prediction
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Glocal Information Bottleneck for Time Series Imputation
- Adapting HFMCA to Graph Data: Self-Supervised Learning for Generalizable fMRI Representations
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
- Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Contrastive Learning for Unpaired Image-to-Image Translation
- Diverse Text-to-Image Generation via Contrastive Noise Optimization
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- DECOR: Deep Embedding Clustering with Orientation Robustness
- Self-Supervised Representation Learning as Mutual Information Maximization
- It Takes Two: Your GRPO Is Secretly DPO
- Feature Identification for Hierarchical Contrastive Learning
- Targeted Supervised Contrastive Learning for Long-Tailed Recognition
- Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
- SimMIM: A Simple Framework for Masked Image Modeling
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
- Towards Intuitive Human-Robot Interaction through Embodied Gesture-Driven Control with Woven Tactile Skins
- Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- PEARL: Performance-Enhanced Aggregated Representation Learning
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- A Family of Kernelized Matrix Costs for Multiple-Output Mixture Neural Networks
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
- Boosting Video Representation Learning with Multi-Faceted Integration
- AraS2P: Arabic Speech-to-Phonemes System
- Personalizing Pre-trained Models
- C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
- Bridging human and machine intelligence: Reverse-engineering radiologist intentions for clinical trust and adoption
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse
- Constructing Contrastive samples via Summarization for Text\n Classification with limited annotations
- Contrastive Video Representation Learning via Adversarial Perturbations
- Learning Generalizable Visual Representations via Interactive Gameplay
- Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
- When Does Self-supervision Improve Few-shot Learning?
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Category Discovery: An Open-World Perspective
- Enriching Knowledge Distillation with Intra-Class Contrastive Learning
- Enhancing Vehicle Detection under Adverse Weather Conditions with Contrastive Learning
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- Generalization in Reinforcement Learning by Soft Data Augmentation
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Efficient Self-supervised Vision Transformers for Representation Learning
- Model Adaptation: Historical Contrastive Learning for Unsupervised Domain Adaptation without Source Data
- Manifold-Aware Diffusion-Augmented Contrastive Learning for Noise-Robust Biosignal Representation
- PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- Bootstrapping User and Item Representations for One-Class Collaborative Filtering
- Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
- Hard Negative Mixing for Contrastive Learning
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- 10 Security and Privacy Problems in Large Foundation Models
- Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- VFM-Loc: Training-Free Cross-View Geo-Localization via Aligning Discriminative Visual Hierarchies
- Temporal Straightening for Latent Planning
- DivCo: Diverse Conditional Image Synthesis via Contrastive Generative Adversarial Network
- Improved Meta-Learning Training for Speaker Verification
- KNN-BERT: Fine-Tuning Pre-Trained Models with KNN Classifier
- Refining Pseudo Labels with Clustering Consensus over Generations for Unsupervised Object Re-identification
- Multimodal Self-Supervised Learning of General Audio Representations
- Understanding Self-supervised Learning with Dual Deep Networks
- TimeSenCLIP: A time series vision–language model for remote sensing
- SSD: A Unified Framework for Self-Supervised Outlier Detection
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- Few-Shot Intent Detection via Contrastive Pre-Training and Fine-Tuning
- BEiT: BERT Pre-Training of Image Transformers
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive Learning
- Coreset selection based on Intra-class diversity
- Global Minimizers of Sigmoid Contrastive Loss
- Hyperbolic Coarse-to-Fine Few-Shot Class-Incremental Learning
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- From Restoration to Reconstruction: Rethinking 3D Gaussian Splatting for Underwater Scenes
- COLA: Context-aware Language-driven Test-time Adaptation
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- Multimodal Medical Image Classification via Synergistic Learning Pre-training
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- Prototypical Contrastive Learning of Unsupervised Representations
- Dynamic Bottleneck for Robust Self-Supervised Exploration
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- Doubly Contrastive Deep Clustering
- UniTac2Pose: A Unified Approach Learned in Simulation for Category-level Visuotactile In-hand Pose Estimation
- Contrastive Learning with Spectrum Information Augmentation in Abnormal Sound Detection
- Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion
- AdaSports-Traj: Role- and Domain-Aware Adaptation for Multi-Agent Trajectory Modeling in Sports
- NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Rethinking "Batch" in BatchNorm
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- IEFS-GMB: Gradient Memory Bank-Guided Feature Selection Based on Information Entropy for EEG Classification of Neurological Disorders
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Training Larger Networks for Deep Reinforcement Learning
- Unsupervised Object-Level Representation Learning from Scene Images
- Interventional Video Grounding with Dual Contrastive Learning
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Masked Feature Modeling Enhances Adaptive Segmentation
- Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Vision-Language Models for Vision Tasks: A Survey
- Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
- Deep Learning-Assisted Detection of Sarcopenia in Cross-Sectional Computed Tomography Imaging
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
- Embodied intelligence via learning and evolution
- MIC: Model-agnostic Integrated Cross-channel Recommenders
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
- Space-Time Correspondence as a Contrastive Random Walk
- Multi-Label Image Classification with Contrastive Learning
- Self-supervised Pretraining of Visual Features in the Wild
- UserBERT: Contrastive User Model Pre-training
- A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
- Text-Based Person Search with Limited Data
- Learning Implicit Sentiment in Aspect-based Sentiment Analysis with Supervised Contrastive Pre-Training
- Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Pre-Trained Image Processing Transformer
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
- Self-supervised learning through the eyes of a child
- Time-Series Representation Learning via Temporal and Contextual Contrasting
- Momentum Contrastive Autoencoder: Using Contrastive Learning for Latent Space Distribution Matching in WAE
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- MCML: A Novel Memory-based Contrastive Meta-Learning Method for Few Shot Slot Tagging
- Dynamic Convolution for 3D Point Cloud Instance Segmentation
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- Improving Audio Event Recognition with Consistency Regularization
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- Preserving Domain Generalization in Fine-Tuning via Joint Parameter Selection
- Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey
- Contrastive Learning with Stronger Augmentations
- Semantic Concentration for Self-Supervised Dense Representations Learning
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Emerging Properties in Self-Supervised Vision Transformers
- Exploring Simple Siamese Representation Learning
- Self-Supervised Graph Learning with Proximity-based Views and Channel Contrast
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Self-Supervised Training Enhances Online Continual Learning
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- RINO: Renormalization Group Invariance with No Labels
- Contrastive Predictive Coding for Anomaly Detection
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic Segmentation
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- Tac2Pose: Tactile object pose estimation from the first touch
- Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Contrastive Learning for Recommender System
- Rethinking Supervised Pre-training for Better Downstream Transferring
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
- Accurate medium-range global weather forecasting with 3D neural networks
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- PLanTS: Periodicity-aware Latent-state Representation Learning for Multivariate Time Series
- USCL: Pretraining Deep Ultrasound Image Diagnosis Model through Video Contrastive Representation Learning
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
- On Fast Adversarial Robustness Adaptation in Model-Agnostic Meta-Learning
- Weakly-Supervised Learning of Dense Functional Correspondences
- The IDLAB VoxCeleb Speaker Recognition Challenge 2020 System Description
- LogME: Practical Assessment of Pre-trained Models for Transfer Learning
- DEMI: Discriminative Estimator of Mutual Information
- Contrastive Neural Processes for Self-Supervised Learning
- Self-training for Few-shot Transfer Across Extreme Task Differences
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- Self-Damaging Contrastive Learning
- CLEAR: Contrastive Learning for Sentence Representation
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- Training GANs with Stronger Augmentations via Contrastive Discriminator
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Contrastive Model Inversion for Data-Free Knowledge Distillation
- Neural Scene Designer: Self-Styled Semantic Image Manipulation
- Self-Supervised Learning for Gastritis Detection with Gastric X-ray Images
- Decomposing and Revising What Language Models Generate
- Bag of Tricks and A Strong baseline for Image Copy Detection
- Partial success in closing the gap between human and machine vision
- HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- CoMET: A Contrastive-Masked Brain Foundation Model for Universal EEG Representation
- Self-Guided Contrastive Learning for BERT Sentence Representations
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Learning from Silence and Noise for Visual Sound Source Localization
- HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
- Representation Learning with Adaptive Superpixel Coding
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- What makes instance discrimination good for transfer learning?
- Spatiotemporal Contrastive Video Representation Learning
- Self-supervised structured object representation learning
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Video Representation Learning with Visual Tempo Consistency
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Self-Supervised Learning with Data Augmentations Provably Isolates\n Content from Style
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- ModAn-MulSupCon: Modality-and Anatomy-Aware Multi-Label Supervised Contrastive Pretraining for Medical Imaging
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
- SoundCLR: Contrastive Learning of Representations For Improved Environmental Sound Classification
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast
- Center-Oriented Prototype Contrastive Clustering
- Behavior From the Void: Unsupervised Active Pre-Training
- Learning ECG Representations via Poly-Window Contrastive Learning
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- Self-Ensembling Contrastive Learning for Semi-Supervised Medical Image Segmentation
- Supporting Clustering with Contrastive Learning
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- A Generalized Learning Framework for Self-Supervised Contrastive Learning
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Multi-view Clustering via Bi-level Decoupling and Consistency Learning
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- Provably Improved Context-Based Offline Meta-RL with Attention and Contrastive Learning
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- CRoC: Context Refactoring Contrast for Graph Anomaly Detection with Limited Supervision
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Click-through Rate Prediction with Auto-Quantized Contrastive Learning
- Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
- Contrastive Weight Regularization for Large Minibatch SGD
- Borrowing From the Future: Enhancing Early Risk Assessment through Contrastive Learning
- Forgery Guided Learning Strategy with Dual Perception Network for Deepfake Cross-domain Detection
- Few Shot Learning With No Labels
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- MCLPD:Multi-view Contrastive Learning for EEG-based PD Detection Across Datasets
- Prototypical Cross-domain Self-supervised Learning for Few-shot Unsupervised Domain Adaptation
- Learning Dense Representations of Phrases at Scale
- IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
- SCALP -- Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
- Multi-Sample based Contrastive Loss for Top-k Recommendation
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Propagation Tree Is Not Deep: Adaptive Graph Contrastive Learning Approach for Rumor Detection
- MISS: Multi-Interest Self-Supervised Learning Framework for Click-Through Rate Prediction
- High-parameter spatial multi-omics through histology-anchored integration
- Artificial intelligence in mitotic checkpoint modeling: transforming our understanding of cellular division through machine learning and predictive biology
- Group-aware Contrastive Regression for Action Quality Assessment
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- CIMON: Towards High-quality Hash Codes
- Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- RRRA: Resampling and Reranking through a Retriever Adapter
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- End-to-end One-shot Human Parsing
- TopKD: Top-scaled Knowledge Distillation
- Bridging Simulation and Experiment: A Self-Supervised Domain Adaptation Framework for Concrete Damage Classification
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- Learning Robust Intervention Representations with Delta Embeddings
- Benchmarking Foundation Models for Mitotic Figure Classification
- Graph Representation Learning with Massive Unlabeled Data for Rumor Detection
- Bootstrap Deep Spectral Clustering with Optimal Transport
- Decoupled Contrastive Learning for Federated Learning
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Scenario-Agnostic Deep-Learning-Based Localization with Contrastive Self-Supervised Pre-training
- Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Semi-Supervised Dual-Threshold Contrastive Learning for Ultrasound Image Classification and Segmentation
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
- MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial Geometry
- Multi-Operator Few-Shot Learning for Generalization Across PDE Families
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- SPENCER: Self-Adaptive Model Distillation for Efficient Code Retrieval
- MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
- Improve Retinal Artery/Vein Classification via Channel Couplin
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Learning Deep Representation with Energy-Based Self-Expressiveness for Subspace Clustering
- AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
Related