MedCLIP: Contrastive Learning from Unpaired Medical Images and Text
2022/10/18 by Zifeng Wang, Wang, Zifeng, Zhenbang Wu +5 · 85 citations
Computer Science · Medicine · #COVID-19 diagnosis using AI #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2210.10163
openalex publication_date 2022/10/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability and supports zero-shot prediction. However, medical image-text datasets are orders of magnitude below the general images and captions from the internet. Moreover, previous methods encounter many false negatives, i.e., images and reports from separate patients probably carry the same semantics but are wrongly treated as negatives. In this paper, we decouple images and texts for multimodal contrastive learning thus scaling the usable training data in a combinatorial magnitude with low cost. We also propose to replace the InfoNCE loss with semantic matching loss based on medical knowledge to eliminate false negatives in contrastive learning. We prove that MedCLIP is a simple yet effective framework: it outperforms state-of-the-art methods on zero-shot prediction, supervised classification, and image-text retrieval. Surprisingly, we observe that with only 20K pre-training data, MedCLIP wins over the state-of-the-art method (using around 200K data). Our code is available at https://github.com/RyanWangZf/MedCLIP.
Cited by
- How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
- NEURO-GUARD: Neuro-Symbolic Generalization and Unbiased Adaptive Routing for Diagnostics -- Explainable Medical AI
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
- Improvise, Adapt, Overcome -- Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- Toward Content-based Indexing and Retrieval of Head and Neck CT with Abscess Segmentation
- Provenance-Driven Reliable Semantic Medical Image Vector Reconstruction via Lightweight Blockchain-Verified Latent Fingerprints
- Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models
- PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation
- Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting
- Uni-Hema: Unified Model for Digital Hematopathology
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
- MedPEFT-CL: Dual-Phase Parameter-Efficient Continual Learning with Medical Semantic Adapter and Bidirectional Memory Consolidation
- NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
- Boosting Medical Visual Understanding From Multi-Granular Language Learning
- D4C: Data-Free Quantization for Contrastive Language-Image Pre-training Models
- DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging
- Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- MAFM3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
- Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography
- T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
- A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models
- MedSAE: Dissecting MedCLIP Representations with Sparse Autoencoders
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
- XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography
- FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning
- Hyperparameter Optimization and Reproducibility in Deep Learning Model Training
- Comprehensive language-image pre-training for 3D medical image understanding
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
- Learning from All: Concept Alignment for Autonomous Distillation from Multiple Drifting MLLMs
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- A Multimodal LLM Approach for Visual Question Answering on Multiparametric 3D Brain MRI
- EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks
- ProbMed: A Probabilistic Framework for Medical Multimodal Binding
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification
- MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness
- Calibration-Aware Prompt Learning for Medical Vision-Language Models
- Exploring the Capabilities of LLM Encoders for Image-Text Retrieval in Chest X-rays
- Multi Anatomy X-Ray Foundation Model
- MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder
- Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance
- GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training
- Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease
- FediLoRA: Heterogeneous LoRA for Federated Multimodal Fine-tuning under Missing Modalities
- Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
- MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction
- Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
- Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Multi-Sequence Parotid Gland Lesion Segmentation via Expert Text-Guided Segment Anything Model
- AMRG: Extend Vision Language Models for Automatic Mammography Report Generation
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
- Benchmarking Foundation Models for Mitotic Figure Classification
- GRIT: Graph-Regularized Logit Refinement for Zero-shot Cell Type Annotation
- Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
- Fairness and Robustness of CLIP-Based Models for Chest X-rays
Related