Improved Baselines with Momentum Contrastive Learning
2020/03/09 by Xinlei Chen, Haoqi Fan, Chen, Xinlei +5 · 413 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Speech Recognition and Synthesis #cs.CV
paper · pdf · doi:10.48550/arxiv.2003.04297
Tech report, 2 pages + references
arxiv created 2020/03/09 · arxiv updated 2020/03/10
Abstract
Contrastive unsupervised learning has recently shown encouraging progress, e.g., in Momentum Contrast (MoCo) and SimCLR. In this note, we verify the effectiveness of two of SimCLR's design improvements by implementing them in the MoCo framework. With simple modifications to MoCo---namely, using an MLP projection head and more data augmentation---we establish stronger baselines that outperform SimCLR and do not require large training batches. We hope this will make state-of-the-art unsupervised learning research more accessible. Code will be made public.
Citations
Cited by
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- Robustifying pathology foundation models via fine-tuning
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- In Pursuit of Pixel Supervision for Visual Pre-training
- M-GRPO: Stabilizing Self-Supervised Reinforcement Learning for Large Language Models with Momentum-Anchored Policy Optimization
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- Super-Selfish: Self-Supervised Learning on Images with PyTorch
- AP-10K: A Benchmark for Animal Pose Estimation in the Wild
- Momentum Contrast for Unsupervised Visual Representation Learning
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Learning Transferable Visual Models From Natural Language Supervision
- Semi-supervised Learning for Dense Object Detection in Retail Scenes
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- Complex QA and language models hybrid architectures, Survey
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Transformers in Vision: A Survey
- On-target Adaptation
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- Contrastive Learning with Hard Negative Samples
- A Survey on Green Deep Learning
- Are we done with ImageNet?
- A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning
- Explicitly Modeling the Discriminability for Instance-Aware Visual Object Tracking
- Cross-domain Contrastive Learning for Unsupervised Domain Adaptation
- Self-supervised Co-training for Video Representation Learning
- Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning
- Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- Large-scale modality-invariant foundation models for brain MRI analysis: Application to lesion segmentation
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- A Comprehensive Study of Deep Video Action Recognition
- Temperature as Uncertainty in Contrastive Learning
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- Exponential Moving Average Normalization for Self-supervised and Semi-supervised Learning
- Self-Supervision with Superpixels: Training Few-shot Medical Image Segmentation without Annotation
- Weakly Supervised Semantic Segmentation by Pixel-to-Prototype Contrast
- SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
- Data-Efficient Instance Generation from Instance Discrimination
- What Makes for Good Views for Contrastive Learning?
- Parametric Contrastive Learning
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- Sensor Data Augmentation by Resampling for Contrastive Learning in Human Activity Recognition
- Self-Supervised Learning with Swin Transformers
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- Pri3D: Can 3D Priors Help 2D Representation Learning?
- Unsupervised Embedding Learning from Uncertainty Momentum Modeling
- Train a One-Million-Way Instance Classifier for Unsupervised Visual Representation Learning
- Do Different Tracking Tasks Require Different Appearance Models?
- Atlas Based Representation and Metric Learning on Manifolds
- Self-supervised similarity search for large scientific datasets
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Unsupervised Local Discrimination for Medical Images
- MIO : Mutual Information Optimization using Self-Supervised Binary Contrastive Learning
- Self-Supervision Closes the Gap Between Weak and Strong Supervision in Histology
- Self-Supervised Multi-View Learning via Auto-Encoding 3D Transformations
- Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift
- On Mutual Information in Contrastive Learning for Visual Representations
- 8-bit Optimizers via Block-wise Quantization
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identification
- Adaptive unified contrastive learning with graph-based feature aggregator for imbalanced medical image classification
- Improve Unsupervised Pretraining for Few-label Transfer
- Exploring Simple Siamese Representation Learning
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Uncovering the structure of clinical EEG signals with self-supervised learning
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- Vector-quantized Image Modeling with Improved VQGAN
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
- GenURL: A General Framework for Unsupervised Representation Learning
- Imbalance-Aware Self-Supervised Learning for 3D Radiomic Representations
- Multi-label Iterated Learning for Image Classification with Label Ambiguity
- Generative Modeling via Drifting
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- MedAug: Contrastive learning leveraging patient metadata improves representations for chest X-ray interpretation
- Using contrastive learning to improve the performance of steganalysis schemes
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
- All you need are a few pixels: semantic segmentation with PixelPick
- Towards the Generalization of Contrastive Self-Supervised Learning
- MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- Latent Diffusion Model without Variational Autoencoder
- Auxiliary Tasks Speed Up Learning PointGoal Navigation
- On Feature Decorrelation in Self-Supervised Learning
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- Contrastive Representations for Label Noise Require Fine-Tuning
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Sparse components distinguish visual pathways & their alignment to neural networks
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- DECOR: Deep Embedding Clustering with Orientation Robustness
- Self-Supervised Representation Learning as Mutual Information Maximization
- Momentum Contrast for Unsupervised Visual Representation Learning
- CODED-SMOOTHING: Coding Theory Helps Generalization
- A Fast Knowledge Distillation Framework for Visual Recognition
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- Personalizing Pre-trained Models
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Generation Properties of Stochastic Interpolation under Finite Training Set
- Efficient Self-supervised Vision Transformers for Representation Learning
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Hard Negative Mixing for Contrastive Learning
- Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- CleftNet: Augmented Deep Learning for Synaptic Cleft Detection from Brain Electron Microscopy
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- Improved Meta-Learning Training for Speaker Verification
- Understanding Self-supervised Learning with Dual Deep Networks
- Self-Supervised Learning with Kernel Dependence Maximization
- BEiT: BERT Pre-Training of Image Transformers
- LoCo: Local Contrastive Representation Learning
- Towards noise robust trigger-word detection with contrastive learning pre-task for fast on-boarding of new trigger-words
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
- Prototypical Contrastive Learning of Unsupervised Representations
- Dynamic Bottleneck for Robust Self-Supervised Exploration
- The SAGES Critical View of Safety Challenge: A Global Benchmark for AI-Assisted Surgical Quality Assessment
- CoUn: Empowering Machine Unlearning via Contrastive Learning
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- Unsupervised Object-Level Representation Learning from Scene Images
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- MIC: Model-agnostic Integrated Cross-channel Recommenders
- Representation Learning via Invariant Causal Mechanisms
- Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images
- Bootstrap your own latent: A new approach to self-supervised Learning
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Self-supervised learning through the eyes of a child
- HyperInverter: Improving StyleGAN Inversion via Hypernetwork
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- Dynamic Convolution for 3D Point Cloud Instance Segmentation
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Contrastive Learning with Stronger Augmentations
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Exploring Simple Siamese Representation Learning
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Self-Supervised Training Enhances Online Continual Learning
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Rethinking Supervised Pre-training for Better Downstream Transferring
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action Segmentation
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- USCL: Pretraining Deep Ultrasound Image Diagnosis Model through Video Contrastive Representation Learning
- Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation
- LogME: Practical Assessment of Pre-trained Models for Transfer Learning
- Self-Damaging Contrastive Learning
- Semi-Supervised Contrastive Learning with Generalized Contrastive Loss and Its Application to Speaker Recognition
- Multimodal Contrastive Training for Visual Representation Learning
- Training GANs with Stronger Augmentations via Contrastive Discriminator
- Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
- Partial success in closing the gap between human and machine vision
- Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Understanding the Behaviour of Contrastive Loss
- Representation Learning with Adaptive Superpixel Coding
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Conditional Negative Sampling for Contrastive Learning of Visual Representations
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- What makes instance discrimination good for transfer learning?
- Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Video Representation Learning with Visual Tempo Consistency
- Whitening for Self-Supervised Representation Learning
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Flatness-aware Curriculum Learning via Adversarial Difficulty
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast
- HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- Supporting Clustering with Contrastive Learning
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- A Generalized Learning Framework for Self-Supervised Contrastive Learning
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose
- Few Shot Learning With No Labels
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- Prototypical Cross-domain Self-supervised Learning for Few-shot Unsupervised Domain Adaptation
- SCALP -- Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
- Contrast R-CNN for Continual Learning in Object Detection
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- Benchmarking Foundation Models for Mitotic Figure Classification
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- Scenario-Agnostic Deep-Learning-Based Localization with Contrastive Self-Supervised Pre-training
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Consistency Regularization for Deep Face Anti-Spoofing
- Unsupervised Visual Representation Learning by Online Constrained K-Means
- A Simple Framework for Uncertainty in Contrastive Learning
- To Label or Not to Label: PALM -- A Predictive Model for Evaluating Sample Efficiency in Active Learning Models
- Improving Joint Embedding Predictive Architecture with Diffusion Noise
- Center-wise Local Image Mixture For Contrastive Representation Learning
- Progressive Stage-wise Learning for Unsupervised Feature Representation Enhancement
- Beneficial Perturbations Network for Defending Adversarial Examples
- MatSSL: Robust Self-Supervised Representation Learning for Metallographic Image Segmentation
- See through Gradients: Image Batch Recovery via GradInversion
- PreViTS: Contrastive Pretraining with Video Tracking Supervision
- MoPro: Webly Supervised Learning with Momentum Prototypes
- Dense Semantic Contrast for Self-Supervised Visual Representation Learning
- Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images
- LEAD: Exploring Logit Space Evolution for Model Selection
- Are Fewer Labels Possible for Few-shot Learning?
- Open-Set Recognition: a Good Closed-Set Classifier is All You Need?
- Attribute Guided Sparse Tensor-Based Model for Person Re-Identification
- Improving Contrastive Learning by Visualizing Feature Transformation
- Image Generators are Generalist Vision Learners
- DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
- Self-EMD: Self-Supervised Object Detection without ImageNet
- Gradient Regularized Contrastive Learning for Continual Domain Adaptation
- ReSSL: Relational Self-Supervised Learning with Weak Augmentation
- 3D Human Action Representation Learning via Cross-View Consistency Pursuit
- MixSiam: A Mixture-based Approach to Self-supervised Representation Learning
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- Knowledge accumulating: The general pattern of learning
- Do We Really Need to Learn Representations from In-domain Data for Outlier Detection?
- Hyperbolic Deep Learning for Foundation Models: A Survey
- Exploring Active Learning for Semiconductor Defect Segmentation
- Principled Multimodal Representation Learning
- Multi-dataset Pretraining: A Unified Model for Semantic Segmentation
- Emerging Properties in Self-Supervised Vision Transformers
- Cluster Contrast for Unsupervised Visual Representation Learning
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- BYOL works even without batch statistics
- CoDiM: Learning with Noisy Labels via Contrastive Semi-Supervised Learning
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- One Million Scenes for Autonomous Driving: ONCE Dataset
- Adaptive Hierarchical Similarity Metric Learning with Noisy Labels
- Self-supervised Visual Attribute Learning for Fashion Compatibility
- Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing
- A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- Robust Contrastive Learning Using Negative Samples with Diminished Semantics
- HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning
- Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- Video Contrastive Learning with Global Context
- Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- Visual Instance-aware Prompt Tuning
- Divergence-Based Similarity Function for Multi-View Contrastive Learning
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- SEED: Self-supervised Distillation For Visual Representation
- Langevin Cooling for Domain Translation
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Integrating Categorical Semantics into Unsupervised Domain Translation
- ST++: Make Self-training Work Better for Semi-supervised Semantic Segmentation
- Compressive Visual Representations
- Cross-Domain Sentiment Classification with In-Domain Contrastive Learning
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- Self-Supervised Pre-Training for Transformer-Based Person Re-Identification
- Landsat-Bench: Datasets and Benchmarks for Landsat Foundation Models
- Self-Supervised Contrastive Learning for Multi-Label Images
- Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
- Cross-Modal Contrastive Learning for Text-to-Image Generation
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Fighting Fire with Fire: Contrastive Debiasing without Bias-free Data via Generative Bias-transformation
- OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical Vision
- Multiple Object Stitching for Unsupervised Representation Learning
- DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
- Self-supervised Learning for Semi-supervised Temporal Language Grounding
- Joint Generative and Contrastive Learning for Unsupervised Person Re-identification
- Spatial-Temporal Pre-Training for Embryo Viability Prediction Using Time-Lapse Videos
- Unsupervised Pre-training for Person Re-identification
- Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
- Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
- Learning Temporally-Consistent Representations for Data-Efficient Reinforcement Learning
- Contrastive Self-Supervised Learning As Neural Manifold Packing
- TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Self Supervision to Distillation for Long-Tailed Visual Recognition
- Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports
- HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks
- Run Away From your Teacher: Understanding BYOL by a Novel Self-Supervised Approach
- When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
- Homography augumented momentum constrastive learning for SAR image retrieval
- MoCo-CXR: MoCo Pretraining Improves Representation and Transferability of Chest X-ray Models
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
- Multi-source Few-shot Domain Adaptation
- FROST: Faster and more Robust One-shot Semi-supervised Training
- Deep Clustering by Semantic Contrastive Learning
- Contrastive Attention for Automatic Chest X-ray Report Generation
- Self-supervised Neural Architecture Search
- What Should Not Be Contrastive in Contrastive Learning
- Improving Transformation Invariance in Contrastive Representation Learning
- Self-Supervised Features Improve Open-World Learning
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Semi-supervised Semantic Segmentation with Directional Context-aware Consistency
- Unleashing the Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-Identification
- How PARTs assemble into wholes: Learning the relative composition of images
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
- Positional Contrastive Learning for Volumetric Medical Image Segmentation
- RadBERT-CL: Factually-Aware Contrastive Learning For Radiology Report Classification
- Online Unsupervised Learning of Visual Representations and Categories
- Jigsaw Clustering for Unsupervised Visual Representation Learning
- Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation
- Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
- To Trust Or Not To Trust Your Vision-Language Model's Prediction
- Few-Shot Segmentation with Global and Local Contrastive Learning
- An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
- Long-Short Temporal Contrastive Learning of Video Transformers
- MixCo: Mix-up Contrastive Learning for Visual Representation
- Object Concepts Emerge from Motion
- Multi-instance Learning as Downstream Task of Self-Supervised Learning-based Pre-trained Model
- Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
- Parametric Instance Classification for Unsupervised Visual Feature Learning
- AmorLIP: Efficient Language-Image Pretraining via Amortization
- Generative AI and foundation models in medical image
- Simple Contrastive Representation Adversarial Learning for NLP Tasks
- Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering
- Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
- Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data
- IDE-Net: Interactive Driving Event and Pattern Extraction from Human Data
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- An Empirical Study and Analysis on Open-Set Semi-Supervised Learning
- Self-supervised Pre-training with Hard Examples Improves Visual Representations
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Momentum2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
- What Is Considered Complete for Visual Recognition?
- Language Models That Walk the Talk: A Framework for Formal Fairness Certificates
- Self-Supervised Learning for Image Segmentation: A Comprehensive Survey
- Semi-Supervised Learning with Taxonomic Labels
- AdaDim: Dimensionality Adaptation for SSL Representational Dynamics
- Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
- Transformer-based unsupervised contrastive learning for histopathological image classification
- S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning
- Mitigating Sampling Bias and Improving Robustness in Active Learning
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Object-Centric Representations Improve Policy Generalization in Robot Manipulation
- Hybrid Discriminative-Generative Training via Contrastive Learning
- A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
- Contrastive Active Inference
- Improving Few-Shot Learning with Auxiliary Self-Supervised Pretext Tasks
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- The Power of Contrast for Feature Learning: A Theoretical Analysis
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
- 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
- MST: Masked Self-Supervised Transformer for Visual Representation
- Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
- Self-Supervised Learning by Estimating Twin Class Distributions
- VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
- Semi-TCL: Semi-Supervised Track Contrastive Representation Learning
- Unsupervised Degradation Representation Learning for Blind Super-Resolution
- Vision Transformers Need More Than Registers
- When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
- Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
- Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
- Benchmarking Transferability: A Framework for Fair and Robust Evaluation
- Estimating Galactic Distances From Images Using Self-supervised Representation Learning
- DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
- Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
- Class-Conditional Distribution Balancing for Group Robust Classification
- SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
- Self-supervised representation learning from 12-lead ECG data
- Don't miss the Mismatch: Investigating the Objective Function Mismatch for Unsupervised Representation Learning
- Self-supervised Cross-silo Federated Neural Architecture Search
- Distilling Localization for Self-Supervised Representation Learning
- I-Con: A Unifying Framework for Representation Learning
- Disentangling Long and Short-Term Interests for Recommendation
- Active Learning with a Noisy Annotator
- Variational Self-Supervised Learning
- Unsupervised Object Detection with LiDAR Clues
- Self-Supervised Multisensor Change Detection
- A Unified Mixture-View Framework for Unsupervised Representation Learning
- Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
- Adapting Vision Foundation Models with Cascaded Semantics
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- Efficient Generative Model Training via Embedded Representation Warmup
- Evolved Hierarchical Masking for Self-Supervised Learning
- Self-Bootstrapping for Versatile Test-Time Adaptation
- Impact of Language Guidance: A Reproducibility Study
- SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
- Joint Learning of Neural Transfer and Architecture Adaptation for Image Recognition
Related