Supervised Contrastive Learning
2020/04/23 by Prannay Khosla, Khosla, Prannay, Piotr Teterwak +15 · 220 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2004.11362
openalex publication_date 2020/04/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Contrastive learning applied to self-supervised representation learning has seen a resurgence in recent years, leading to state of the art performance in the unsupervised training of deep image models. Modern batch contrastive approaches subsume or significantly outperform traditional contrastive losses such as triplet, max-margin and the N-pairs loss. In this work, we extend the self-supervised batch contrastive approach to the fully-supervised setting, allowing us to effectively leverage label information. Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes. We analyze two possible versions of the supervised contrastive (SupCon) loss, identifying the best-performing formulation of the loss. On ResNet-200, we achieve top-1 accuracy of 81.4% on the ImageNet dataset, which is 0.8% above the best number reported for this architecture. We show consistent outperformance over cross-entropy on other datasets and two ResNet variants. The loss shows benefits for robustness to natural corruptions and is more stable to hyperparameter settings such as optimizers and data augmentations. Our loss function is simple to implement, and reference TensorFlow code is released at https://t.ly/supcon.
Citations
Cited by
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
- Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation
- Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
- Controlling Embedding Spaces with Text-Conditioned Transformations
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- STAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting
- Text2Graph VPR: A Text-to-Graph Expert System for Explainable Place Recognition in Changing Environments
- Modality-Dependent Memory Mechanisms in Cross-Modal Neuromorphic Computing
- The Urysohn Ladder: Recursive Metric Contraction for Scalable Continual Learning
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- MACL: Multi-Label Adaptive Contrastive Learning Loss for Remote Sensing Image Retrieval
- Open Ad-hoc Categorization with Contextualized Feature Learning
- SCS-SupCon: Sigmoid-based Common and Style Supervised Contrastive Learning with Adaptive Decision Boundaries
- Automated Motion Artifact Check for MRI (AutoMAC-MRI): An Interpretable Framework for Motion Artifact Detection and Severity Assessment
- CF-Net: A Cross-Feature Reconstruction Network for High-Accuracy 1-Bit Target Classification
- A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments
- Enhancing Semi-Supervised Multi-View Graph Convolutional Networks via Supervised Contrastive Learning and Self-Training
- Predictive Sample Assignment for Semantically Coherent Out-of-Distribution Detection
- Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
- Supervised Contrastive Frame Aggregation for Video Representation Learning
- AMBER: An Adaptive Multimodal Mask Transformer for Beam Prediction with Missing Modalities
- LogICL: Distilling LLM Reasoning to Bridge the Semantic Gap in Cross-Domain Log Anomaly Detection
- Contrastive Learning for Semi-Supervised Deep Regression with Generalized Ordinal Rankings from Spectral Seriation
- CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification
- ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
- Personalized Image Descriptions from Attention Sequences
- DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
- FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
- Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
- Domain Feature Collapse: Implications for Out-of-Distribution Detection and Solutions
- Story2MIDI: Emotionally Aligned Music Generation from Text
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- MANTA: Physics-Informed Generalized Underwater Object Tracking
- Robust HRRP Recognition under Interrupted Sampling Repeater Jamming using a Prior Jamming Information-Guided Network
- Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
- Semi-Supervised Contrastive Learning with Orthonormal Prototypes
- Graph Contrastive Learning via Spectral Graph Alignment
- PISA: Prioritized Invariant Subgraph Aggregation
- MRI-Based Brain Age Estimation with Supervised Contrastive Learning of Continuous Representation
- Dynamic Stratified Contrastive Learning with Upstream Augmentation for MILP Branching
- RLM: A Vision-Language Model Approach for Radar Scene Understanding
- Point-Supervised Facial Expression Spotting with Gaussian-Based Instance-Adaptive Intensity Modeling
- Late-decoupled 3D Hierarchical Semantic Segmentation with Semantic Prototype Discrimination based Bi-branch Supervision
- Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
- Boosting Medical Visual Understanding From Multi-Granular Language Learning
- MAPROC at AHaSIS Shared Task: Few-Shot and Sentence Transformer for Sentiment Analysis of Arabic Hotel Reviews
- UniSER: A Foundation Model for Unified Soft Effects Removal
- FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding
- SLAM-AGS: Slide-Label Aware Multi-Task Pretraining Using Adaptive Gradient Surgery in Computational Cytology
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive Projection
- Quantum Machine Learning via Contrastive Training
- An Evaluation of Representation Learning Methods in Particle Physics Foundation Models
- Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- CITADEL: A Semi-Supervised Active Learning Framework for Malware Detection Under Continuous Distribution Drift
- HEDN: A Hard-Easy Dual Network with Source Reliability Assessment for Cross-Subject EEG Emotion Recognition
- Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment
- Climbing the label tree: Hierarchy-preserving contrastive learning for medical imaging
- Image-Intrinsic Priors for Integrated Circuit Defect Detection and Novel Class Discovery via Self-Supervised Learning
- MiRAGE: Misconception Detection with Retrieval-Guided Multi-Stage Reasoning and Ensemble Fusion
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
- Efficient Online Continual Learning in Sensor-Based Human Activity Recognition
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Metadata-Aligned 3D MRI Representations for Contrast Understanding and Quality Control
- Functional embeddings enable Aggregation of multi-area SEEG recordings over subjects and sessions
- ANCHOR: Integrating Adversarial Training with Hard-mined Supervised Contrastive Learning for Robust Representation Learning
- Mitigating Semantic Collapse in Partially Relevant Video Retrieval
- When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
- A Theory of Contrastive Learning with Natural Images
- Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
- Torus embeddings
- Robust and Generalizable Atrial Fibrillation Detection from ECG Using Time-Frequency Fusion and Supervised Contrastive Learning
- DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis
- A Dual-Branch CNN for Robust Detection of AI-Generated Facial Forgeries
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- Multi-dataset Joint Pre-training of Emotional EEG Enables Generalizable Affective Computing
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Transformed Multi-view 3D Shape Features with Contrastive Learning
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- AWSPNet: Attention-based Dual-Tree Wavelet Scattering Prototypical Network for MIMO Radar Target Recognition and Jamming Suppression
- CARE: Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams
- Evaluating protein binding interfaces with PUMBA
- Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- SDGraph: Multi-Level Sketch Representation Learning by Sparse-Dense Graph Architecture
- Simulation-Based Pretraining and Domain Adaptation for Astronomical Time Series with Minimal Labeled Data
- Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Source-Free Object Detection with Detection Transformer
- rareboost3d: a synthetic lidar dataset with enhanced rare classes
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- FairContrast: Enhancing Fairness through Contrastive learning and Customized Augmenting Methods on Tabular Data
- Generative design of synthetic gene circuits for functional and evolutionary properties
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Function regression using the forward forward training and inferring paradigm
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Contrastive Representation Regularization for Vision-Language-Action Models
- Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph Prediction
- Learning Representations Through Contrastive Neural Model Checking
- Contrastive-SDE: Guiding Stochastic Differential Equations with Contrastive Learning for Unpaired Image-to-Image Translation
- Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
- Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- LLM Routing with Dueling Feedback
- Feature Identification for Hierarchical Contrastive Learning
- VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
- Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Contrastive Diffusion Guidance for Spatial Inverse Problems
- OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- Generalized Category Discovery in Hyperspectral Images via Prototype Subspace Modeling
- Online Specific Emitter Identification via Collision-Alleviated Signal Hash
- Contrastive Learning Enhances Language Model Based Cell Embeddings for Low-Sample Single Cell Transcriptomics
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Graph Your Own Prompt
- Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- FlowXpert: Context-Aware Flow Embedding for Enhanced Traffic Detection in IoT Network
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
- What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
- When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment
- LLM2Vec-Gen: Generative Embeddings from Large Language Models
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- Towards noise robust trigger-word detection with contrastive learning\n pre-task for fast on-boarding of new trigger-words
- M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition
- Robustness Feature Adapter for Efficient Adversarial Training
- Contrastive Learning for Robust Android Malware Familial Classification
- MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion
- Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- Long-Tailed Out-of-Distribution Detection with Refined Separate Class Learning
- Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports
- CoUn: Empowering Machine Unlearning via Contrastive Learning
- DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis
- EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
- Global Pre-fixing, Local Adjusting: A Simple yet Effective Contrastive Strategy for Continual Learning
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- PhenoGnet: A Graph-Based Contrastive Learning Framework for Disease Similarity Prediction
- Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
- NORA: A Nephrology-Oriented Representation Learning Approach Towards Chronic Kidney Disease Classification
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- Learning Representations in Video Game Agents with Supervised Contrastive Imitation Learning
- Detecting Multilevel Manipulation from Limit Order Book via Cascaded Contrastive Representation Learning
- VH-Diffuser: Variable Horizon Diffusion Planner for Time-Aware Goal-Conditioned Trajectory Planning
- Promoting Shape Bias in CNNs: Frequency-Based and Contrastive Regularization for Corruption Robustness
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
- Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey
- TUNI: Unifying Pre-training and Fine-tuning with Modality-Aware Mutual Learning and Rectification for RGB-T Semantic Segmentation
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Advancing Few-Shot Pediatric Arrhythmia Classification with a Novel Contrastive Loss and Multimodal Learning
- Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- Tackling the Noisy Elephant in the Room: Label Noise-robust Out-of-Distribution Detection via Loss Correction and Low-rank Decomposition
- Contrastive Anatomy-Contrast Disentanglement: A Domain-General MRI Harmonization Method
- CAME-AB: Cross-Modality Attention with Mixture-of-Experts for Antibody Binding Site Prediction
- Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
- Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
- FedQuad: Federated Stochastic Quadruplet Learning to Mitigate Data Heterogeneity
- On Episodes, Prototypical Networks, and Few-shot Learning
- MorphGen: Morphology-Guided Representation Learning for Robust Single-Domain Generalization in Histopathological Cancer Classification
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Multimodal Fusion Refiner Networks
- Foundation Model-Driven Classification of Atypical Mitotic Figures with Domain-Aware Training Strategies
- Generalizable Object Re-Identification via Visual In-Context Prompting
- "Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
- A Survey of Affective Recommender Systems: Modeling Attitudes, Emotions, and Moods for Personalization
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- NM-Hebb: Coupling Local Hebbian Plasticity with Metric Learning for More Accurate and Interpretable CNNs
- Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
- ModAn-MulSupCon: Modality-and Anatomy-Aware Multi-Label Supervised Contrastive Pretraining for Medical Imaging
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
- SoundCLR: Contrastive Learning of Representations For Improved Environmental Sound Classification
- Comprehensively stratifying MCIs into distinct risk subtypes based on brain imaging genetics fusion learning
- Evolution Is All You Need: Phylogenetic Augmentation for Contrastive Learning
- Paired-Sampling Contrastive Framework for Joint Physical-Digital Face Attack Detection
- Deep Learning for Taxol Exposure Analysis: A New Cell Image Dataset and Attention-Based Baseline Model
- Learning Point Cloud Representations with Pose Continuity for Depth-Based Category-Level 6D Object Pose Estimation
- Shift Detection and Adaptation for Network Intrusion Detection
- Supporting Clustering with Contrastive Learning
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- Fracture Detection and Localisation in Wrist and Hand Radiographs using Detection Transformer Variants
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- Multi-Domain Supervised Contrastive Learning for UAV Radio-Frequency Open-Set Recognition
- Few-shot Class-incremental Fault Diagnosis by Preserving Class-Agnostic Knowledge with Dual-Granularity Representations
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- Multi-level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- Domain Generalization of Pathological Image Segmentation by Patch-Level and WSI-Level Contrastive Learning
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- Graph-based Robot Localization Using a Graph Neural Network with a Floor Camera and a Feature Rich Industrial Floor
- Attribute Guidance With Inherent Pseudo-label For Occluded Person Re-identification
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- Learning Robust Intervention Representations with Delta Embeddings
- WSS-CL: Weight Saliency Soft-Guided Contrastive Learning for Efficient Machine Unlearning Image Classification
- Dynamic User-controllable Privacy-preserving Few-shot Sensing Framework
- Pseudo-label Induced Subspace Representation Learning for Robust Out-of-Distribution Detection
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
- TrajSurv: Learning Continuous Latent Trajectories from Electronic Health Records for Trustworthy Survival Prediction
- From LLMs to Edge: Parameter-Efficient Fine-Tuning on Edge Devices
- From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
Related