Momentum Contrast for Unsupervised Visual Representation Learning
2020/06/01 by Kaiming He, Haoqi Fan, Yuxin Wu +2 · 1,191 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications
paper · doi:10.1109/cvpr42600.2020.00975
openalex publication_date 2020/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.
Citations
Cited by
- Self-Supervised Graph Co-Training for Session-based Recommendation
- ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- Pri3D: Can 3D Priors Help 2D Representation Learning?
- X-model: Improving Data Efficiency in Deep Learning with A Minimax Model
- Deep Long-Tailed Learning: A Survey
- Do Different Tracking Tasks Require Different Appearance Models?
- Unsupervised Natural Language Inference via Decoupled Multimodal Contrastive Learning
- PGL: Prior-Guided Local Self-supervised Learning for 3D Medical Image Segmentation
- Self-Supervised Multi-View Learning via Auto-Encoding 3D Transformations
- Mine Your Own vieW: Self-Supervised Learning Through Across-Sample\n Prediction
- Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift
- Self-supervised Heterogeneous Graph Neural Network with Co-contrastive Learning
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identification
- CLCC: Contrastive Learning for Color Constancy
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Deep semi-supervised learning for medical image segmentation: A review
- 3D Human Pose, Shape and Texture from Low-Resolution Images and Videos
- Structural-Spectral Graph Convolution with Evidential Edge Learning for Hyperspectral Image Clustering
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
- Deep Learning with Label Differential Privacy
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Uncovering the structure of clinical EEG signals with self-supervised learning
- Pairwise Supervised Contrastive Learning of Sentence Representations
- Self-supervised Contrastive Video-Speech Representation Learning for Ultrasound
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- High-Performance Large-Scale Image Recognition Without Normalization
- Sparsity-Probe: Analysis tool for Deep Learning Models
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- Vector-quantized Image Modeling with Improved VQGAN
- Video Understanding as Machine Translation
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Deep Learning for Person Re-Identification: A Survey and Outlook
- Contrastive Feature Loss for Image Prediction
- Contrastive Semi-Supervised Learning for 2D Medical Image Segmentation
- Dual Contrastive Learning for Unsupervised Image-to-Image Translation
- Pre-training Molecular Graph Representation with 3D Geometry
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
- GenURL: A General Framework for Unsupervised Representation Learning
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- Trustworthy AI: From Principles to Practices
- Fundamental Limits and Tradeoffs in Invariant Representation Learning
- Automatic Shortcut Removal for Self-Supervised Representation Learning
- Imbalance-Aware Self-Supervised Learning for 3D Radiomic Representations
- CLDA: Contrastive Learning for Semi-Supervised Domain Adaptation
- Learning to Prompt for Vision-Language Models
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Exploring Data-Efficient 3D Scene Understanding with Contrastive Scene Contexts
- Cycle-Contrast for Self-Supervised Video Representation Learning
- RL makes MLLMs see better than SFT
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- Eigenfunction Extraction for Ordered Representation Learning
- Graph Neural Networks: Methods, Applications, and Opportunities
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Image Quality Assessment using Contrastive Learning
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
- Domain Adaptive Semantic Segmentation with Self-Supervised Depth Estimation
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- Mutual Information guided Visual Contrastive Learning
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
- Deep Graph Contrastive Representation Learning
- Unsupervised Part Discovery from Contrastive Reconstruction
- Contrastive Learning for Many-to-many Multilingual Neural Machine Translation
- Resounding Acoustic Fields with Reciprocity
- Exploring Conditions for Diffusion models in Robotic Control
- MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
- Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Towards the Generalization of Contrastive Self-Supervised Learning
- CIL: Contrastive Instance Learning Framework for Distantly Supervised Relation Extraction
- ε-Seg: Sparsely Supervised Semantic Segmentation of Microscopy Data
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Self-Supervised Longitudinal Neighbourhood Embedding
- Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
- Token-Level Inference-Time Alignment for Vision-Language Models
- MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- Confidence-Weighted Semi-Supervised Learning for Skin Lesion Segmentation Using Hybrid CNN-Transformer Networks
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Geospatial Machine Learning Libraries
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Latent Diffusion Model without Variational Autoencoder
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Morphology-Aware Prognostic model for Five-Year Survival Prediction in Colorectal Cancer from H&E Whole Slide Images
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- ViTacGen: Robotic Pushing with Vision-to-Touch Generation
- Progressive Cluster Purification for Unsupervised Feature Learning
- Rethinking Graph Domain Adaptation: A Spectral Contrastive Perspective
- Universal Image Restoration Pre-training via Masked Degradation Classification
- Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation
- On Feature Decorrelation in Self-Supervised Learning
- Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
- FedHUG: Federated Heterogeneous Unsupervised Generalization for Remote Physiological Measurements
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- Diffusion Transformers with Representation Autoencoders
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Are Labels Necessary for Neural Architecture Search?
- Feature Stylization and Domain-aware Contrastive Learning for Domain Generalization
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Contrastive Dimension Reduction: A Systematic Review
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Scaling Language-Centric Omnimodal Representation Learning
- Contrastive Noise-Guided Invertible Network for Image Steganography
- Rectifying the Shortcut Learning of Background for Few-Shot Learning
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- Redundancy as a Structural Information Principle for Learning and Generalization
- A Joint Learning Approach to Hardware Caching and Prefetching
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- Decoupling Representation Learning from Reinforcement Learning
- Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
- Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss
- Astronomia ex machina: a history, primer and outlook on neural networks in astronomy
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- 3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Vision Language Models: A Survey of 26K Papers
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- Conditional Alignment and Uniformity for Contrastive Learning with Continuous Proxy Labels
- Contrastive Representations for Label Noise Require Fine-Tuning
- A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Batch Curation for Unsupervised Contrastive Representation Learning
- Contrastive Self-Supervised Learning at the Edge: An Energy Perspective
- Label-Efficient Multi-Task Segmentation using Contrastive Learning
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- DADO: A Depth-Attention framework for Object Discovery
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
- SSL-SE-EEG: A Framework for Robust Learning from Unlabeled EEG Data with Self-Supervised Learning and Squeeze-Excitation Networks
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- Conditional Representation Learning for Customized Tasks
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Contrastive Representation Regularization for Vision-Language-Action Models
- Cross-Batch Negative Sampling for Training Two-Tower Recommenders
- Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph Prediction
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Glocal Information Bottleneck for Time Series Imputation
- Adapting HFMCA to Graph Data: Self-Supervised Learning for Generalizable fMRI Representations
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
- Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Contrastive Learning for Unpaired Image-to-Image Translation
- Diverse Text-to-Image Generation via Contrastive Noise Optimization
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- DECOR: Deep Embedding Clustering with Orientation Robustness
- Self-Supervised Representation Learning as Mutual Information Maximization
- It Takes Two: Your GRPO Is Secretly DPO
- Feature Identification for Hierarchical Contrastive Learning
- Targeted Supervised Contrastive Learning for Long-Tailed Recognition
- Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
- SimMIM: A Simple Framework for Masked Image Modeling
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
- Towards Intuitive Human-Robot Interaction through Embodied Gesture-Driven Control with Woven Tactile Skins
- Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- PEARL: Performance-Enhanced Aggregated Representation Learning
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- A Family of Kernelized Matrix Costs for Multiple-Output Mixture Neural Networks
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
- Boosting Video Representation Learning with Multi-Faceted Integration
- AraS2P: Arabic Speech-to-Phonemes System
- Personalizing Pre-trained Models
- C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
- Bridging human and machine intelligence: Reverse-engineering radiologist intentions for clinical trust and adoption
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse
- Constructing Contrastive samples via Summarization for Text Classification with limited annotations
- Contrastive Video Representation Learning via Adversarial Perturbations
- Learning Generalizable Visual Representations via Interactive Gameplay
- Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
- When Does Self-supervision Improve Few-shot Learning?
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Category Discovery: An Open-World Perspective
- Enriching Knowledge Distillation with Intra-Class Contrastive Learning
- Enhancing Vehicle Detection under Adverse Weather Conditions with Contrastive Learning
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- Generalization in Reinforcement Learning by Soft Data Augmentation
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Efficient Self-supervised Vision Transformers for Representation Learning
- Model Adaptation: Historical Contrastive Learning for Unsupervised Domain Adaptation without Source Data
- Manifold-Aware Diffusion-Augmented Contrastive Learning for Noise-Robust Biosignal Representation
- PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
- Bootstrapping User and Item Representations for One-Class Collaborative Filtering
- Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
- Hard Negative Mixing for Contrastive Learning
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- 10 Security and Privacy Problems in Large Foundation Models
- Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- Learning Curves for Analysis of Deep Networks
- VFM-Loc: Training-Free Cross-View Geo-Localization via Aligning Discriminative Visual Hierarchies
- Temporal Straightening for Latent Planning
- DivCo: Diverse Conditional Image Synthesis via Contrastive Generative Adversarial Network
- Improved Meta-Learning Training for Speaker Verification
- KNN-BERT: Fine-Tuning Pre-Trained Models with KNN Classifier
- Refining Pseudo Labels with Clustering Consensus over Generations for Unsupervised Object Re-identification
- Multimodal Self-Supervised Learning of General Audio Representations
- Understanding Self-supervised Learning with Dual Deep Networks
- TimeSenCLIP: A time series vision–language model for remote sensing
- SSD: A Unified Framework for Self-Supervised Outlier Detection
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- Few-Shot Intent Detection via Contrastive Pre-Training and Fine-Tuning
- BEiT: BERT Pre-Training of Image Transformers
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive Learning
- Coreset selection based on Intra-class diversity
- Global Minimizers of Sigmoid Contrastive Loss
- Hyperbolic Coarse-to-Fine Few-Shot Class-Incremental Learning
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- From Restoration to Reconstruction: Rethinking 3D Gaussian Splatting for Underwater Scenes
- COLA: Context-aware Language-driven Test-time Adaptation
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- Multimodal Medical Image Classification via Synergistic Learning Pre-training
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- Prototypical Contrastive Learning of Unsupervised Representations
- Dynamic Bottleneck for Robust Self-Supervised Exploration
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- Doubly Contrastive Deep Clustering
- UniTac2Pose: A Unified Approach Learned in Simulation for Category-level Visuotactile In-hand Pose Estimation
- Contrastive Learning with Spectrum Information Augmentation in Abnormal Sound Detection
- Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion
- AdaSports-Traj: Role- and Domain-Aware Adaptation for Multi-Agent Trajectory Modeling in Sports
- NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Rethinking "Batch" in BatchNorm
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- IEFS-GMB: Gradient Memory Bank-Guided Feature Selection Based on Information Entropy for EEG Classification of Neurological Disorders
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Training Larger Networks for Deep Reinforcement Learning
- Unsupervised Object-Level Representation Learning from Scene Images
- Interventional Video Grounding with Dual Contrastive Learning
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Masked Feature Modeling Enhances Adaptive Segmentation
- Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Vision-Language Models for Vision Tasks: A Survey
- Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
- Deep Learning-Assisted Detection of Sarcopenia in Cross-Sectional Computed Tomography Imaging
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
- Embodied Intelligence via Learning and Evolution
- MIC: Model-agnostic Integrated Cross-channel Recommenders
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
- Space-Time Correspondence as a Contrastive Random Walk
- Multi-Label Image Classification with Contrastive Learning
- Self-supervised Pretraining of Visual Features in the Wild
- UserBERT: Contrastive User Model Pre-training
- A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
- Text-Based Person Search with Limited Data
- Learning Implicit Sentiment in Aspect-based Sentiment Analysis with Supervised Contrastive Pre-Training
- Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Pre-Trained Image Processing Transformer
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
- Self-supervised learning through the eyes of a child
- Time-Series Representation Learning via Temporal and Contextual Contrasting
- Momentum Contrastive Autoencoder: Using Contrastive Learning for Latent Space Distribution Matching in WAE
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- MCML: A Novel Memory-based Contrastive Meta-Learning Method for Few Shot Slot Tagging
- Dynamic Convolution for 3D Point Cloud Instance Segmentation
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- Improving Audio Event Recognition with Consistency Regularization
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- Preserving Domain Generalization in Fine-Tuning via Joint Parameter Selection
- Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey
- Contrastive Learning with Stronger Augmentations
- Semantic Concentration for Self-Supervised Dense Representations Learning
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Exploring Simple Siamese Representation Learning
- Self-Supervised Graph Learning with Proximity-based Views and Channel Contrast
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Self-Supervised Training Enhances Online Continual Learning
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- RINO: Renormalization Group Invariance with No Labels
- Contrastive Predictive Coding for Anomaly Detection
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic Segmentation
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- Tac2Pose: Tactile object pose estimation from the first touch
- Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Contrastive Learning for Recommender System
- Rethinking Supervised Pre-training for Better Downstream Transferring
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
- Accurate medium-range global weather forecasting with 3D neural networks
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- PLanTS: Periodicity-aware Latent-state Representation Learning for Multivariate Time Series
- USCL: Pretraining Deep Ultrasound Image Diagnosis Model through Video Contrastive Representation Learning
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
- On Fast Adversarial Robustness Adaptation in Model-Agnostic Meta-Learning
- Weakly-Supervised Learning of Dense Functional Correspondences
- The IDLAB VoxCeleb Speaker Recognition Challenge 2020 System Description
- LogME: Practical Assessment of Pre-trained Models for Transfer Learning
- DEMI: Discriminative Estimator of Mutual Information
- Contrastive Neural Processes for Self-Supervised Learning
- Self-training for Few-shot Transfer Across Extreme Task Differences
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- Self-Damaging Contrastive Learning
- CLEAR: Contrastive Learning for Sentence Representation
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- Training GANs with Stronger Augmentations via Contrastive Discriminator
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Contrastive Model Inversion for Data-Free Knowledge Distillation
- Neural Scene Designer: Self-Styled Semantic Image Manipulation
- Self-Supervised Learning for Gastritis Detection with Gastric X-ray Images
- Decomposing and Revising What Language Models Generate
- Bag of Tricks and A Strong baseline for Image Copy Detection
- Partial success in closing the gap between human and machine vision
- HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- CoMET: A Contrastive-Masked Brain Foundation Model for Universal EEG Representation
- Self-Guided Contrastive Learning for BERT Sentence Representations
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Learning from Silence and Noise for Visual Sound Source Localization
- HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
- Representation Learning with Adaptive Superpixel Coding
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- What makes instance discrimination good for transfer learning?
- Spatiotemporal Contrastive Video Representation Learning
- Self-supervised structured object representation learning
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Video Representation Learning with Visual Tempo Consistency
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- ModAn-MulSupCon: Modality-and Anatomy-Aware Multi-Label Supervised Contrastive Pretraining for Medical Imaging
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
- SoundCLR: Contrastive Learning of Representations For Improved Environmental Sound Classification
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast
- Center-Oriented Prototype Contrastive Clustering
- Behavior From the Void: Unsupervised Active Pre-Training
- Learning ECG Representations via Poly-Window Contrastive Learning
- Semantic-Aware Generation for Self-Supervised Visual Representation Learning
- MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- Self-Ensembling Contrastive Learning for Semi-Supervised Medical Image Segmentation
- Supporting Clustering with Contrastive Learning
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- A Generalized Learning Framework for Self-Supervised Contrastive Learning
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Multi-view Clustering via Bi-level Decoupling and Consistency Learning
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- Provably Improved Context-Based Offline Meta-RL with Attention and Contrastive Learning
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- CRoC: Context Refactoring Contrast for Graph Anomaly Detection with Limited Supervision
- DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort
- Dual-view Molecule Pre-training
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Click-through Rate Prediction with Auto-Quantized Contrastive Learning
- Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation
- Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
- Contrastive Weight Regularization for Large Minibatch SGD
- Borrowing From the Future: Enhancing Early Risk Assessment through Contrastive Learning
- Forgery Guided Learning Strategy with Dual Perception Network for Deepfake Cross-domain Detection
- Few Shot Learning With No Labels
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- MCLPD:Multi-view Contrastive Learning for EEG-based PD Detection Across Datasets
- Prototypical Cross-domain Self-supervised Learning for Few-shot Unsupervised Domain Adaptation
- Learning Dense Representations of Phrases at Scale
- IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
- SCALP -- Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
- Multi-Sample based Contrastive Loss for Top-k Recommendation
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Propagation Tree Is Not Deep: Adaptive Graph Contrastive Learning Approach for Rumor Detection
- MISS: Multi-Interest Self-Supervised Learning Framework for Click-Through Rate Prediction
- High-parameter spatial multi-omics through histology-anchored integration
- Artificial intelligence in mitotic checkpoint modeling: transforming our understanding of cellular division through machine learning and predictive biology
- Group-aware Contrastive Regression for Action Quality Assessment
- Learning to Align Sequential Actions in the Wild
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- CIMON: Towards High-quality Hash Codes
- Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- RRRA: Resampling and Reranking through a Retriever Adapter
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- End-to-end One-shot Human Parsing
- TopKD: Top-scaled Knowledge Distillation
- Bridging Simulation and Experiment: A Self-Supervised Domain Adaptation Framework for Concrete Damage Classification
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- Learning Robust Intervention Representations with Delta Embeddings
- Benchmarking Foundation Models for Mitotic Figure Classification
- Graph Representation Learning with Massive Unlabeled Data for Rumor Detection
- Bootstrap Deep Spectral Clustering with Optimal Transport
- Decoupled Contrastive Learning for Federated Learning
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Scenario-Agnostic Deep-Learning-Based Localization with Contrastive Self-Supervised Pre-training
- Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Semi-Supervised Dual-Threshold Contrastive Learning for Ultrasound Image Classification and Segmentation
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
- MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial Geometry
- Multi-Operator Few-Shot Learning for Generalization Across PDE Families
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- SPENCER: Self-Adaptive Model Distillation for Efficient Code Retrieval
- MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
- Improve Retinal Artery/Vein Classification via Channel Couplin
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Learning Deep Representation with Energy-Based Self-Expressiveness for Subspace Clustering
- AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
- Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
- Consistency Regularization for Deep Face Anti-Spoofing
- Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
- Revisiting Catastrophic Forgetting in Class Incremental Learning
- Ensemble Foreground Management for Unsupervised Object Discovery
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Unsupervised Visual Representation Learning by Online Constrained K-Means
- Learning and Evaluating Representations for Deep One-class Classification
- A Simple Framework for Uncertainty in Contrastive Learning
- Style-Aware Blending and Prototype-Based Cross-Contrast Consistency for Semi-Supervised Medical Image Segmentation
- Online Adversarial Purification based on Self-Supervision
- Measuring Generalization with Optimal Transport
- Rethinking Pre-training and Self-training
- One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation
- A mini-batch training strategy for deep subspace clustering networks
- Improving Joint Embedding Predictive Architecture with Diffusion Noise
- Latent Denoising Makes Good Tokenizers
- FBNetV3: Joint Architecture-Recipe Search using Predictor Pretraining
- Self-supervised human mobility learning for next location prediction and trajectory classification
- UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
- Center-wise Local Image Mixture For Contrastive Representation Learning
- Progressive Stage-wise Learning for Unsupervised Feature Representation Enhancement
- Longitudinal Correlation Analysis for Decoding Multi-Modal Brain Development
- Self-Guided Masked Autoencoder
- Iterative Graph Self-Distillation
- SpecBPP: A Self-Supervised Learning Approach for Hyperspectral Representation and Soil Organic Carbon Estimation
- MC-SSL0.0: Towards Multi-Concept Self-Supervised Learning
- Pneumonia Detection on Chest X-ray using Radiomic Features and Contrastive Learning
- Vectorization and Rasterization: Self-Supervised Learning for Sketch and Handwriting
- Evaluating Modules in Graph Contrastive Learning
- CLEAR: Unlearning Spurious Style-Content Associations with Contrastive LEarning with Anti-contrastive Regularization
- Towards Zero-shot Commonsense Reasoning with Self-supervised Refinement of Language Models
- Quantifying and Mitigating Privacy Risks of Contrastive Learning
- ROBAD: Robust Adversary-aware Local-Global Attended Bad Actor Detection Sequential Model
- See through Gradients: Image Batch Recovery via GradInversion
- Hierarchical Cross-modal Prompt Learning for Vision-Language Models
- User Invariant Preference Learning for Multi-Behavior Recommendation
- Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation
- Cluster Contrast for Unsupervised Person Re-Identification
- PreViTS: Contrastive Pretraining with Video Tracking Supervision
- Discriminative-Generative Representation Learning for One-Class Anomaly Detection
- Self-Distilled Self-Supervised Representation Learning
- Modeling optical imaging pipeline and learning contrastive-based representation for hybrid-corrupted image restoration
- Dense Semantic Contrast for Self-Supervised Visual Representation Learning
- Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images
- LEAD: Exploring Logit Space Evolution for Model Selection
- Are Fewer Labels Possible for Few-shot Learning?
- Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
- OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
- A Comprehensive Survey on Source-Free Domain Adaptation
- UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
- The History of Speech Recognition to the Year 2030
- Improving Contrastive Learning by Visualizing Feature Transformation
- Sentence Semantic Regression for Text Generation
- Image Generators are Generalist Vision Learners
- Support-set bottlenecks for video-text representation learning
- Revisiting Graph Contrastive Learning on Anomaly Detection: A Structural Imbalance Perspective
- DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
- Socially-Aware Self-Supervised Tri-Training for Recommendation
- Self-EMD: Self-Supervised Object Detection without ImageNet
- Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization
- Self-Supervised Sketch-to-Image Synthesis
- Enhanced Seq2Seq Autoencoder via Contrastive Learning for Abstractive Text Summarization
- About contrastive unsupervised representation learning for classification and its convergence
- Gradient Regularized Contrastive Learning for Continual Domain Adaptation
- The 2021 Image Similarity Dataset and Challenge
- Inter-intra Variant Dual Representations forSelf-supervised Video Recognition
- EnAET: A Self-Trained framework for Semi-Supervised and Supervised Learning with Ensemble Transformations
- 3D Human Action Representation Learning via Cross-View Consistency Pursuit
- On Inductive Biases for Machine Learning in Data Constrained Settings
- MixSiam: A Mixture-based Approach to Self-supervised Representation Learning
- Transferrable Contrastive Learning for Visual Domain Adaptation
- Unsupervised Point Cloud Pre-Training via Occlusion Completion
- Proceedings of the First Workshop on Weakly Supervised Learning (WeaSuL)
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Unsupervised Representation Learning via Neural Activation Coding
- Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification
- Do We Really Need to Learn Representations from In-domain Data for Outlier Detection?
- CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
- C3RL: Rethinking the Combination of Channel-independence and Channel-mixing from Representation Learning
- Exploring Active Learning for Semiconductor Defect Segmentation
- Principled Multimodal Representation Learning
- Multi-Level Transfer Learning from Near-Field to Far-Field Speaker Verification
- CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation
- Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning
- CM-UNet: A Self-Supervised Learning-Based Model for Coronary Artery Segmentation in X-Ray Angiography
- Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
- MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks
- Multi-dataset Pretraining: A Unified Model for Semantic Segmentation
- Adapt, But Don't Forget: Fine-Tuning and Contrastive Routing for Lane Detection under Distribution Shift
- Towards channel foundation models (CFMs): Motivations, methodologies and opportunities
- Leveraging Language Prior for Infrared Small Target Detection
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
- Emerging Properties in Self-Supervised Vision Transformers
- GLOMIA-Pro: A Generalizable Longitudinal Medical Image Analysis Framework for Disease Progression Prediction
- Cluster Contrast for Unsupervised Visual Representation Learning
- Domain-Aware Augmentations for Unsupervised Online General Continual Learning
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- BYOL works even without batch statistics
- Relating by Contrasting: A Data-efficient Framework for Multimodal Generative Models
- The Impact of Spatiotemporal Augmentations on Self-Supervised Audiovisual Representation Learning
- S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning
- CoDiM: Learning with Noisy Labels via Contrastive Semi-Supervised Learning
- Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- One Million Scenes for Autonomous Driving: ONCE Dataset
- ActiveMatch: End-to-end Semi-supervised Active Representation Learning
- Calibrating and Improving Graph Contrastive Learning
- Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person Search
- MT3: Meta Test-Time Training for Self-Supervised Test-Time Adaption
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Mixing-AdaSIN: Constructing a De-biased Dataset using Adaptive Structural Instance Normalization and Texture Mixing
- Pre-Trained Models: Past, Present and Future
- Towards Domain-Agnostic Contrastive Learning
- Multi-Format Contrastive Learning of Audio Representations
- Diffuse and Disperse: Image Generation with Representation Regularization
- Multimodal Representation Alignment for Cross-modal Information Retrieval
- Neural, Symbolic and Neural-Symbolic Reasoning on Knowledge Graphs
- Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
- Bi-tuning of Pre-trained Representations
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
- Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
- Clustering-Guided Multi-Layer Contrastive Representation Learning for Citrus Disease Classification
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- DepViT-CAD: Deployable Vision Transformer-Based Cancer Diagnosis in Histopathology
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
- Robust Contrastive Learning Using Negative Samples with Diminished Semantics
- Foundation Models in Medical Imaging: A Review and Outlook
- Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- Video Contrastive Learning with Global Context
- Pre-training without Natural Images
- Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
- PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies
- ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning
- Intention-Conditioned Flow Occupancy Models
- Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
- Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model
- PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Self-Supervised Learning at the Edge: The Cost of Labeling
- A Principled Framework for Multi-View Contrastive Learning
- Towards Discriminative Representation Learning for Unsupervised Person Re-identification
- CL-Polyp: A Contrastive Learning-Enhanced Network for Accurate Polyp Segmentation
- GreenHyperSpectra: A multi-source hyperspectral dataset for global vegetation trait prediction
- Ambiguity-aware Point Cloud Segmentation by Adaptive Margin Contrastive Learning
- Divergence-Based Similarity Function for Multi-View Contrastive Learning
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
- DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- SEED: Self-supervised Distillation For Visual Representation
- UnrealPerson: An Adaptive Pipeline towards Costless Person Re-identification
- AAG: Self-Supervised Representation Learning by Auxiliary Augmentation with GNT-Xent Loss
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
- Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
- COIN: Contrastive Identifier Network for Breast Mass Diagnosis in Mammography
- Semi-weakly Supervised Contrastive Representation Learning for Retinal Fundus Images
- PointGAC: Geometric-Aware Codebook for Masked Point Cloud Modeling
- Information-Guided Diffusion Sampling for Dataset Distillation
- GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
- Representation learning with a transformer by contrastive learning for money laundering detection
- Langevin Cooling for Domain Translation
- Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval
- Improving Context-Based Meta-Reinforcement Learning with Self-Supervised Trajectory Contrastive Learning
- Physics-Aligned Self-Supervised Learning for Scientific Imaging
- MMOC: Self-Supervised EEG Emotion Recognition Framework with Multi-Model Online Collaboration
- Graph-MLP: Node Classification without Message Passing in Graph
- A Note on Connecting Barlow Twins with Negative-Sample-Free Contrastive Learning
- Rectifying Adversarial Sample with Low Entropy Prior for Test-Time Defense
- Contrastive Learning with Temporal Correlated Medical Images: A Case Study using Lung Segmentation in Chest X-Rays
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
- Integrating Categorical Semantics into Unsupervised Domain Translation
- PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
- Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning
- Prototypical Contrast and Reverse Prediction: Unsupervised Skeleton Based Action Recognition
- MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping
- A Self-Supervised Framework for Function Learning and Extrapolation
- AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction
- Robust brain age estimation from structural MRI with contrastive learning
- Compressive Visual Representations
- Hierarchical Self-Supervised Learning for Medical Image Segmentation Based on Multi-Domain Data Aggregation
- Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
- Medical Image Segmentation with Limited Supervision: A Review of Deep Network Models
- m-RevNet: Deep Reversible Neural Networks with Momentum
- Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment
- CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation
- Unsupervised Representation Learning for Binary Networks by Joint Classifier Learning
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
- VOCAL: Visual Odometry via ContrAstive Learning
- Embedding-based Retrieval in Multimodal Content Moderation
- When Test-Time Adaptation Meets Self-Supervised Models
- A Theoretical Formulation on the Use of Multiple Positive Views in Contrastive Learning
- From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
- FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation
- High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
- DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
- Hierarchical Corpus-View-Category Refinement for Carotid Plaque Risk Grading in Ultrasound
- Self-Supervised Contrastive Learning for Multi-Label Images
- Pareto Self-Supervised Training for Few-Shot Learning
- Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
- Attention to the Burstiness in Visual Prompt Tuning!
- LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning
- Cross-Modal Contrastive Learning for Text-to-Image Generation
- Neural Feature Search for RGB-Infrared Person Re-Identification
- Multilingual BERT Post-Pretraining Alignment
- Multi-View Contrastive Learning for Robust Domain Adaptation in Medical Time Series Analysis
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- FedSC: Federated Learning with Semantic-Aware Collaboration
- Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- ConCAD: Contrastive Learning-based Cross Attention for Sleep Apnea Detection
- rQdia: Regularizing Q-Value Distributions With Image Augmentation
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical Vision
- ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations
- IQFM A Wireless Foundational Model for I/Q Streams in AI-Native 6G
- Multiple Object Stitching for Unsupervised Representation Learning
- Iterative Quantum Feature Maps
- Self-supervised Temporal Discriminative Learning for Video Representation Learning
- Contrastive Cross-Modal Learning for Infusing Chest X-ray Knowledge into ECGs
- Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
- AquaCluster: Using Satellite Images And Self-supervised Machine Learning Networks To Detect Water Hidden Under Vegetation
- Resampling Augmentation for Time Series Contrastive Learning: Application to Remote Sensing
- Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework
- DIP: Unsupervised Dense In-Context Post-training of Visual Representations
- Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review
- Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
- Variational Supervised Contrastive Learning
- Benchmarking Foundation Models and Parameter-Efficient Fine-Tuning for Prognosis Prediction in Medical Imaging
- DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
- A Survey on False Information Detection: From A Perspective of Propagation on Social Networks
- Joint Generative and Contrastive Learning for Unsupervised Person Re-identification
- PhysiX: A Foundation Model for Physics Simulations
- Spatial-Temporal Pre-Training for Embryo Viability Prediction Using Time-Lapse Videos
- Self-supervised Feature Extraction for Enhanced Ball Detection on Soccer Robots
- Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
- Unsupervised Pre-training for Person Re-identification
- A Simple Contrastive Framework Of Item Tokenization For Generative Recommendation
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- Constrained Contrastive Distribution Learning for Unsupervised Anomaly Detection and Localisation in Medical Images
- Self-supervised Text-independent Speaker Verification using Prototypical Momentum Contrastive Learning
- CRIA: A Cross-View Interaction and Instance-Adapted Pre-training Framework for Generalizable EEG Representations
- Bridging Brain with Foundation Models through Self-Supervised Learning
- Cluster Analysis with Deep Embeddings and Contrastive Learning
- Learning of Inter-Label Geometric Relationships Using Self-Supervised Learning: Application To Gleason Grade Segmentation
- Reliable Few-shot Learning under Dual Noises
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer
- Reinforcement Learning with Augmented Data
- Pretraining Representations for Data-Efficient Reinforcement Learning
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- TrajDiff: Diffusion Bridge Network with Semantic Alignment for Trajectory Similarity Computation
- Play to Generalize: Learning to Reason Through Game Play
- Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
- Semi-supervised Long-tailed Recognition using Alternate Sampling
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
- Contrastive Self-Supervised Learning As Neural Manifold Packing
- TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast
- SeqPE: Transformer with Sequential Position Encoding
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
- SimTriplet: Simple Triplet Representation Learning with a Single GPU
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Unsupervised Contrastive Learning Using Out-Of-Distribution Data for Long-Tailed Dataset
- Long-Tailed Learning for Generalized Category Discovery
- Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
- Contrastive Unsupervised Learning for Speech Emotion Recognition
- Information fusion strategy integrating pre-trained language model and contrastive learning for materials knowledge mining
- Self Supervision to Distillation for Long-Tailed Visual Recognition
- Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
- DiffFuSR: Super-Resolution of all Sentinel-2 Multispectral Bands using Diffusion Models
- Hard-Soft Pseudo Labels Guided Semi-Supervised Learning for Point Cloud Classification
- Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports
- UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning
- Visual Pre-Training on Unlabeled Images using Reinforcement Learning
- Simpler, Faster, Stronger: Breaking The log-K Curse On Contrastive Learners With FlatNCE
- Structure-Aware Automatic Channel Pruning by Searching with Graph Embedding
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation
- Rethinking Random Masking in Self-Distillation on ViT
- Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks
- Run Away From your Teacher: Understanding BYOL by a Novel Self-Supervised Approach
- Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection
- The Role of Global Labels in Few-Shot Classification and How to Infer Them
- Rethinking Graph Contrastive Learning through Relative Similarity Preservation
- Representation Learning for Sequence Data with Deep Autoencoding Predictive Components
- TreeGAN: Incorporating Class Hierarchy into Image Generation
- Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning
- Blind Image Super-Resolution via Contrastive Representation Learning
- Learning to Weight Parameters for Training Data Attribution
- Distribution Estimation to Automate Transformation Policies for Self-Supervision
- When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
- Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
- Homography augumented momentum constrastive learning for SAR image retrieval
- Multi-axis Attentive Prediction for Sparse EventData: An Application to Crime Prediction
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation
- Contrastive Flow Matching
- Mitigating Degree Bias Adaptively with Hard-to-Learn Nodes in Graph Contrastive Learning
- Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual Distractions
- Tuning the Right Foundation Models is What you Need for Partial Label Learning
- FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
- Mixture-of-Experts Meets In-Context Reinforcement Learning
- VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
- ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning
- Self-labelling via simultaneous clustering and representation learning
- Investigating the Role of Negatives in Contrastive Representation Learning
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- Fine-Grained Image Analysis with Deep Learning: A Survey
- CLIPort: What and Where Pathways for Robotic Manipulation
- DREAM: Visual Decoding from Reversing Human Visual System
- Language-Driven Image Style Transfer
- Biomedical Entity Linking with Contrastive Context Matching
- Multi-source Few-shot Domain Adaptation
- WDMamba: When Wavelet Degradation Prior Meets Vision Mamba for Image Dehazing
- Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
- Prune Your Model Before Distill It
- Deep Clustering by Semantic Contrastive Learning
- Meta-Learning to Improve Pre-Training
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation
- Breaking Shortcut: Exploring Fully Convolutional Cycle-Consistency for Video Correspondence Learning
- Contrastive Attention for Automatic Chest X-ray Report Generation
- Self-supervised Neural Architecture Search
- What Should Not Be Contrastive in Contrastive Learning
- Self-Supervised Features Improve Open-World Learning
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Semi-supervised Semantic Segmentation with Directional Context-aware Consistency
- Can Temporal Information Help with Contrastive Self-Supervised Learning?
- eProduct: A Million-Scale Visual Search Benchmark to Address Product Recognition Challenges
- SemiOccam: A Robust Semi-Supervised Image Recognition Network Using Sparse Labels
- Robust Neural Rendering in the Wild with Asymmetric Dual 3D Gaussian Splatting
- Self-Supervised Ranking for Representation Learning
- Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
- Unleashing the Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-Identification
- How PARTs assemble into wholes: Learning the relative composition of images
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
- Joint Generalized Cosine Similarity: A Novel Method for N-Modal Semantic Alignment Based on Contrastive Learning
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
- Illiterate DALL-E Learns to Compose
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- Self-supervised Representation Learning for Evolutionary Neural Architecture Search
- ES-Net: An Efficient Stereo Matching Network
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
- Text Data Augmentation for Deep Learning
- ConMamba: Contrastive Vision Mamba for Plant Disease Detection
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Self-Supervised Spatial Correspondence Across Modalities
- Towards a Hypothesis on Visual Transformation based Self-Supervision
- Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
- Data Augmentation for Object Detection via Differentiable Neural Rendering
- Boosting Few-Shot Classification with View-Learnable Contrastive Learning
- Aligned Contrastive Loss for Long-Tailed Recognition
- Positional Contrastive Learning for Volumetric Medical Image Segmentation
- Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges
- ECP-Mamba: An Efficient Multi-scale Self-supervised Contrastive Learning Method with State Space Model for PolSAR Image Classification
- RadBERT-CL: Factually-Aware Contrastive Learning For Radiology Report Classification
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- Perceptual Inductive Bias Is What You Need Before Contrastive Learning
- Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
- Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
- Adversarial Training with Contrastive Learning in NLP
- Online Unsupervised Learning of Visual Representations and Categories
- JojoSCL: Shrinkage Contrastive Learning for single-cell RNA sequence Clustering
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation Learning
- Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
- Jigsaw Clustering for Unsupervised Visual Representation Learning
- Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
- Unleashing the Power of Intermediate Domains for Mixed Domain Semi-Supervised Medical Image Segmentation
- GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training
- CrossTransformers: spatially-aware few-shot transfer
- A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
- A Tutorial-cum-Survey on Self-Supervised Learning for Wi-Fi Sensing: Trends, Challenges, and Outlook
- Self-supervised feature learning for cardiac Cine MR image reconstruction
- CoSformer: Detecting Co-Salient Object with Transformers
- Patient Domain Supervised Contrastive Learning for Lung Sound Classification Using Mobile Phone
- Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
- Balancing Robustness and Sensitivity using Feature Contrastive Learning
- IDEAL: Independent Domain Embedding Augmentation Learning
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- What is the Right Embedding Space for Contrastive Learning in REC?
- A Survey on Long-Tailed Visual Recognition
- A Comprehensive Survey on Test-Time Adaptation Under Distribution Shifts
- Contrastive Learning and Self-Training for Unsupervised Domain Adaptation in Semantic Segmentation
- Unsupervised Domain Adaptive Learning via Synthetic Data for Person Re-identification
- Few-Shot Segmentation with Global and Local Contrastive Learning
- D2LV: A Data-Driven and Local-Verification Approach for Image Copy Detection
- Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
- Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
- Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
- AgriFM: A Multi-source Temporal Remote Sensing Foundation Model for Crop Mapping
- DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
- Long-Short Temporal Contrastive Learning of Video Transformers
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- Object Concepts Emerge from Motion
- Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
- Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition
- Learning Object-Centric Video Models by Contrasting Sets
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- Dual Contrastive Loss and Attention for GANs
- Enhancing Contrastive Learning-based Electrocardiogram Pretrained Model with Patient Memory Queue
- A Contrastive Learning Foundation Model Based on Perfectly Aligned Sample Pairs for Remote Sensing Images
- Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
- FruitNeRF++: A Generalized Multi-Fruit Counting Method Utilizing Contrastive Learning and Neural Radiance Fields
- Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model
- NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application
- Parametric Instance Classification for Unsupervised Visual Feature\n Learning
- Advancing Video Self-Supervised Learning via Image Foundation Models
- Label Contrastive Coding based Graph Neural Network for Graph Classification
- Anycost GANs for Interactive Image Synthesis and Editing
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
- AmorLIP: Efficient Language-Image Pretraining via Amortization
- FedSKC: Federated Learning with Non-IID Data via Structural Knowledge Collaboration
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
- Dual-Path Stable Soft Prompt Generation for Domain Generalization
- Disentangling Knowledge Representations for Large Language Model Editing
- Generative AI and foundation models in medical image
- Contrastive Learning for Mitochondria Segmentation
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- Aggregative Self-Supervised Feature Learning from a Limited Sample
- Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
- Structured Proxy Features for Multimodal NSCLC Survival Prediction from Pretreatment CT
- Simple Contrastive Representation Adversarial Learning for NLP Tasks
- Robust Robotic Control from Pixels using Contrastive Recurrent State-Space Models
- SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
- Contrastive Learning based Hybrid Networks for Long-Tailed Image Classification
- Accelerating Learned Image Compression Through Modeling Neural Training Dynamics
- Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering
- Towards Compact Single Image Super-Resolution via Contrastive Self-distillation
- Generative Models as a Data Source for Multiview Representation Learning
- Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- 3D Self-Supervised Methods for Medical Imaging
- Intermediate Layers Matter in Momentum Contrastive Self Supervised Learning
- Do DeepFake Attribution Models Generalize?
- An Empirical Study of Graph Contrastive Learning
- SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- Exploiting Sample Uncertainty for Domain Adaptive Person Re-Identification
- Unsupervised Person Re-identification via Simultaneous Clustering and Consistency Learning
- Unsupervised deep learning for grading of age-related macular degeneration using retinal fundus images
- An Empirical Study and Analysis on Open-Set Semi-Supervised Learning
- Self-supervised Pre-training with Hard Examples Improves Visual Representations
- Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
- Understanding and Achieving Efficient Robustness with Adversarial Supervised Contrastive Learning
- VET-DINO: Learning Anatomical Understanding Through Multi-View Distillation in Veterinary Imaging
- gen2seg: Generative Models Enable Generalizable Instance Segmentation
- Learning From Long-Tailed Data With Noisy Labels
- Khan-GCL: Kolmogorov-Arnold Network Based Graph Contrastive Learning with Hard Negatives
- SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification
- Field Matters: A lightweight LLM-enhanced Method for CTR Prediction
- Generative Hierarchical Features from Synthesizing Images
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- Bootstrapping Semantic Segmentation with Regional Contrast
- Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
- Collaborative Unlabeled Data Optimization
- CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding
- Self-Reinforced Graph Contrastive Learning
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
- Free-Form Image Inpainting via Contrastive Attention Network
- Momentum2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
- What Is Considered Complete for Visual Recognition?
- Contrastive Mixup: Self- and Semi-Supervised learning for Tabular Domain
- FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
- Language Models That Walk the Talk: A Framework for Formal Fairness Certificates
- GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
- Weakly Supervised Contrastive Learning for Chest X-Ray Report Generation
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction
- Self-Supervised Learning for Image Segmentation: A Comprehensive Survey
- Simplicity is Key: An Unsupervised Pretraining Approach for Sparse Radio Channels
- Semi-Supervised Learning with Taxonomic Labels
- The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
- Multi-Modal Artificial Intelligence of Embryo Grading and Pregnancy Prediction in Assisted Reproductive Technology: A Review
- Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
- Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
- Unsupervised Document Embedding via Contrastive Augmentation
- Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
- Equally Critical: Samples, Targets, and Their Mappings in Datasets
- NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation
- S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- Temporal Context Aggregation for Video Retrieval with Contrastive Learning
- Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic Learning
- A novel multiple instance learning framework for COVID-19 severity assessment via data augmentation and self-supervised learning
- Pix2seq: A Language Modeling Framework for Object Detection
- Self-supervised Remote Sensing Images Change Detection at Pixel-level
- Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning
- Exploiting Intrinsic Duality for Multi-Hop Question Generation
- Mitigating Sampling Bias and Improving Robustness in Active Learning
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Maximizing Asynchronicity in Event-based Neural Networks
- Self-Supervised Policy Adaptation during Deployment
- Is Supervised Learning Really That Different from Unsupervised?
- CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
- KFCNet: Knowledge Filtering and Contrastive Learning Network for Generative Commonsense Reasoning
- Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
- MPPFND: A Dataset and Analysis of Detecting Fake News with Multi-Platform Propagation
- OpenMatch: Open-set Consistency Regularization for Semi-supervised Learning with Outliers
- Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning
- Memory Augmented Multi-Instance Contrastive Predictive Coding for Sequential Recommendation
- A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
- EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
- ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
- EDBench: Large-Scale Electron Density Data for Molecular Modeling
- VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction
- Domain-Agnostic Clustering with Self-Distillation
- Exploring Task Difficulty for Few-Shot Relation Extraction
- Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
- Contrastive Active Inference
- Integrating Auxiliary Information in Self-supervised Learning
- TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series
- Take More Positives: An Empirical Study of Contrastive Learing in Unsupervised Person Re-Identification
- Improving Few-Shot Learning with Auxiliary Self-Supervised Pretext Tasks
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative Augmentation
- RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation Learning
- Contrastive and Generative Graph Convolutional Networks for Graph-based Semi-Supervised Learning
- The Power of Contrast for Feature Learning: A Theoretical Analysis
- Image Classification Using a Diffusion Model as a Pre-Training Model
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- Deep Robust Clustering by Contrastive Learning
- Feature Representation Transferring to Lightweight Models via Perception Coherence
- A Contrastive Federated Semi-Supervised Learning Intrusion Detection Framework for Internet of Robotic Things
- SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
- PSSL: Self-supervised Learning for Personalized Search with Contrastive Sampling
- Unlocking the Full Potential of Small Data with Diverse Supervision
- UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
- Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
- FG-CLIP: Fine-Grained Visual and Textual Alignment
- How to build the best medical image segmentation algorithm using foundation models: a comprehensive empirical study with Segment Anything Model
- ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications
- Multi-Behavior Hypergraph Contrastive Learning for Session-Based Recommendation
- Rethinking Sampling Strategies for Unsupervised Person Re-identification
- Dynamic Feature Alignment for Semi-supervised Domain Adaptation
- Magnification-independent Histopathological Image Classification with Similarity-based Multi-scale Embeddings
- Self-supervised Representation Learning with Relative Predictive Coding
- 6-DoF Contrastive Grasp Proposal Network
- Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
- Self-Supervision Enhances Instance-based Multiple Instance Learning Methods in Digital Pathology: A Benchmark Study
- PREMISE: Matching-based Prediction for Accurate Review Recommendation
- Energy Aligning for Biased Models
- Self-Supervised Learning by Estimating Twin Class Distributions
- Supervised Momentum Contrastive Learning for Few-Shot Classification
- Code-Enhanced Cross-Perspective Bug Question Retrieval
- Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
- Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- INSID3: Training-Free In-Context Segmentation with DINOv3
- Implicit Neural-Representation Learning for Elastic Deformable-Object Manipulations
- Contrastive Embedding for Generalized Zero-Shot Learning
- Multi-aspect Graph Contrastive Learning for Review-enhanced Recommendation
- Improved Mutual Mean-Teaching for Unsupervised Domain Adaptive Re-ID
- User identification network with contrastive clustering for shared-account recommendation
- Graph Contrastive Pre-training for Effective Theorem Reasoning
- OneFeed: A Unified Generative Framework for Feed Content Enhancement and Query Generation
- Learning Retrospective Knowledge with Reverse Reinforcement Learning
- Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- Color Variants Identification in Fashion e-commerce via Contrastive Self-Supervised Representation Learning
- Semi-TCL: Semi-Supervised Track Contrastive Representation Learning
- Unsupervised Degradation Representation Learning for Blind Super-Resolution
- The Spatially-Correlative Loss for Various Image Translation Tasks
- BERT Bi-modal self-supervised learning for crop classification using Sentinel-2 and Planetscope
- When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
- FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Distilling LLM Feedback for Lean Theorem Proving
- On the Dynamics of Observation and Semantics
- StreamSplit: Continuous Audio Representation Learning via Uncertainty-Guided Adaptive Splitting
- When Does LeJEPA Learn a World Model?
- Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
- Recursive KL Divergence Optimization: A Dynamic Framework for Representation Learning
- Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
- Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
- Style-Adaptive Detection Transformer for Single-Source Domain Generalized Object Detection
- SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
- DeepAndes: A Self-Supervised Vision Foundation Model for Multi-Spectral Remote Sensing Imagery of the Andes
- Triangular Consistency as a Universal Constraint for Learning Optical Flow
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations
- Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
- Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
- Comparison of Image Processing Models in Quark Gluon Jet Classification
- Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning
- Benchmarking Transferability: A Framework for Fair and Robust Evaluation
- High-precision large-aperture single-frame interferometric surface profile measurement method based on deep learning
- MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis
- Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID
- Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation
- CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- ASMa: Asymmetric Spatio-temporal Masking for Skeleton Action Representation Learning
- A Closer Look at Few-Shot Video Classification: A New Baseline and Benchmark
- SleepPriorCL: Contrastive Representation Learning with Prior Knowledge-based Positive Mining and Adaptive Temperature for Sleep Staging
- Inverse Problems Leveraging Pre-trained Contrastive Representations
- Estimating Galactic Distances From Images Using Self-supervised Representation Learning
- Self-Supervised Video Representation Learning by Video Incoherence Detection
- Contrastive Multiview Coding
- Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training
- Learning a Domain-Agnostic Visual Representation for Autonomous Driving via Contrastive Loss
- DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
- Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition
- Earth Embeddings
- Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
- Speaker Verification Under Real Classroom Conditions for English Speech
- Class-Conditional Distribution Balancing for Group Robust Classification
- Contrastive representation enhancement and learning for handwritten mathematical expression recognition
- SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
- Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
- A Genealogy of Foundation Models in Remote Sensing
- NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
- Contrastive Supervised Distillation for Continual Representation Learning
- Training Vision Transformers for Image Retrieval
- XSimGCL: Towards Extremely Simple Graph Contrastive Learning for Recommendation
- Contrastive Semi-supervised Learning for ASR
- Exploring Dense Retrieval for Dialogue Response Selection
- Self-supervised representation learning from 12-lead ECG data
- Towards to Robust and Generalized Medical Image Segmentation Framework
- CSS-LM: A Contrastive Framework for Semi-Supervised Fine-Tuning of Pre-Trained Language Models
- Revisiting Radar Camera Alignment by Contrastive Learning for 3D Object Detection
- Representation Learning via Non-Contrastive Mutual Information
- RPT: Toward Transferable Model on Heterogeneous Researcher Data via Pre-Training
- Don't miss the Mismatch: Investigating the Objective Function Mismatch for Unsupervised Representation Learning
- Learning Domain-Agnostic Visual Representation for Computational Pathology Using Medically-Irrelevant Style Transfer Augmentation
- Self-supervised Cross-silo Federated Neural Architecture Search
- MMHCL: Multi-Modal Hypergraph Contrastive Learning for Recommendation
- Federated EndoViT: Pretraining Vision Transformers via Federated Learning on Endoscopic Image Collections
- Disentangling Long and Short-Term Interests for Recommendation
- Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical Imaging
- Anomaly Detection on Attributed Networks via Contrastive Self-Supervised Learning
- Unsupervised Object Detection with LiDAR Clues
- VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
- Representation Synthesis by Probabilistic Many-Valued Logic Operation in Self-Supervised Learning
- Self-Supervised Change Detection in Multiview Remote Sensing Images
- Integrating Contrastive Learning with Dynamic Models for Reinforcement Learning from Images
- Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length Adjustment
- SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-training
- SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- Decomposition-based multi-scale transformer framework for time series anomaly detection
- DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
- A Unified Mixture-View Framework for Unsupervised Representation Learning
- Unsupervised Deep Representation Learning and Few-Shot Classification of PolSAR Images
- When Contrastive Learning Meets Active Learning: A Novel Graph Active Learning Paradigm with Self-Supervision
- Unsupervised Hashing with Contrastive Information Bottleneck
- Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-Supervised Speaker Verification
- Dual Contradistinctive Generative Autoencoder
- Self-supervised Graph Neural Networks without explicit negative sampling
- Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data for User Perception
- RADLER: Radar Object Detection Leveraging Semantic 3D City Models and Self-Supervised Radar-Image Learning
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- From Local Learning to Global Prediction Through Layered Surprise Cascades
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
- MIEB: Massive Image Embedding Benchmark
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- Efficient Generative Model Training via Embedded Representation Warmup
- Unveiling Contrastive Learning's Capability of Neighborhood Aggregation for Collaborative Filtering
- AimTS: Augmented Series and Image Contrastive Learning for Time Series Classification
- Intent Contrastive Learning for Sequential Recommendation
- Evolved Hierarchical Masking for Self-Supervised Learning
- Boosting multi-demographic federated learning for chest radiograph analysis using general-purpose self-supervised representations
- Impact of Language Guidance: A Reproducibility Study
- Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
- CHIME: A Compressive Framework for Holistic Interest Modeling
- SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic Segmentation
- On the Importance of Conditioning for Privacy-Preserving Data Augmentation
- Joint Learning of Neural Transfer and Architecture Adaptation for Image Recognition
- Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
- Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
- Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
- DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
- MIMRS: A Survey on Masked Image Modeling in Remote Sensing
- RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
- AdaViT: Adaptive Vision Transformer for Flexible Pretrain and Finetune with Variable 3D Medical Image Modalities
- Towards Generalizing Temporal Action Segmentation to Unseen Views
- X-Capture: An Open-Source Portable Device for Multi-Sensory Learning
- Deep Momentum Uncertainty Hashing
- Learning from Streaming Video with Orthogonal Gradients
- Scene-Centric Unsupervised Panoptic Segmentation
- FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
- Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
- MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation
- Vernata: Self-Supervised Learning of LiDAR Point Representations
Related