A data- and compute-efficient chest X-ray foundation model beyond aggressive scaling
2026/02/26 by Chong Wang, Yabin Zhang, Yunhe Gao +9 · 1 voice
Computer Science · #cs.CV
paper · pdf · doi:10.48550/arxiv.2602.22843
Abstract
Foundation models for medical imaging are typically pretrained on increasingly large datasets, following a "scale-at-all-costs" paradigm. However, this strategy faces two critical challenges: large-scale medical datasets often contain substantial redundancy and severe class imbalance that bias representation learning toward over-represented patterns, and indiscriminate training regardless of heterogeneity in data quality incurs considerable computational inefficiency. Here we demonstrate that active, principled data curation during pretraining can serve as a viable, cost-effective alternative to brute-force dataset enlargement. We introduce CheXficient, a chest X-ray (CXR) foundation model that selectively prioritizes informative training samples. CheXficient is pretrained on only 22.7% of 1,235,004 paired CXR images and reports while consuming under 27.3% of the total compute budget, yet achieving comparable or superior performance to its full-data counterpart and other large-scale pretrained models. We assess CheXficient across 20 individual benchmarks spanning 5 task types, including non-adapted off-the-shelf evaluations (zero-shot findings classification and crossmodal retrieval) and adapted downstream tasks (disease prediction, semantic segmentation, and radiology report generation). Further analyses show that CheXficient systematically prioritizes under-represented training samples, improving generalizability on long-tailed or rare conditions. Overall, our work offers practical insights into the data and computation demands for efficient pretraining and downstream adaptation of medical vision-language foundation models.
Citations
- MedGemma Technical Report
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
- ReXGradient-160K: A Large-Scale Publicly Available Dataset of Chest Radiographs with Free-text Reports
- BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
- Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis
- A Multimodal Vision Foundation Model for Clinical Dermatology
- RaTEScore: A Metric for Radiology Report Generation
- A Comprehensive Survey of Foundation Models in Medicine
- MAIRA-2: Grounded Radiology Report Generation
- CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
- MedVersa: A Generalist Foundation Model for Medical Image Interpretation
- Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
- Foundation Model for Advancing Healthcare: Challenges, Opportunities, and Future Directions
- Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
- Mixture of Gaussian-distributed Prototypes with Generative Modelling for Interpretable and Trustworthy Image Recognition
- MedFMC: A Real-world Dataset and Benchmark For Foundation Model Adaptation in Medical Image Classification
- Too Large; Data Reduction for Vision-Language Pre-Training
- DataComp: In search of the next generation of multimodal datasets
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- Interpretable Medical Image Visual Question Answering via Multi-Modal Relationship Graph Learning
- Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing
- Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards
- Beyond neural scaling laws: beating power law scaling via data pruning
- Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt
- Rethinking Semantic Segmentation: A Prototype View
- PediCXR: An open, large-scale chest radiograph dataset for interpretation of common thoracic diseases in children
- On the Opportunities and Risks of Foundation Models
- VinDr-RibCXR: A Benchmark Dataset for Automatic Segmentation and Labeling of Individual Ribs on Chest X-rays
- Learning Transferable Visual Models From Natural Language Supervision
- VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises
- BIMCV COVID-19+: a large annotated dataset of RX and CT images from COVID-19 patients
- Quantifying the Value of Lateral Views in Deep Learning for Chest X-rays
- BERTScore: Evaluating Text Generation with BERT
- Publicly Available Clinical BERT Embeddings
- PadChest: A large chest x-ray image dataset with multi-label annotated reports
- Representation Learning with Contrastive Predictive Coding
- Decoupled Weight Decay Regularization
- ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Sinkhorn Distances: Lightspeed Computation of Optimal Transportation Distances
- A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
- Development of a Digital Image Database for Chest Radiographs With and Without a Lung Nodule
Discussions
Related