vix.ing · top · new · best · stats · spec

Vision-language model-based semantic-guided imaging biomarker for lung nodule malignancy prediction

2025/04/30 by Luoting Zhuang, Zhuang, Luoting, Seyed Mohammad Hossein Tabatabaei +11 · 1 voice
Medicine · #Lung Cancer Diagnosis and Treatment #COVID-19 diagnosis using AI #Radiomics and Machine Learning in Medical Imaging

paper · doi:10.1016/j.jbi.2025.104947

openalex created_date 2025/10/10 · openalex publication_date 2025/10/27 · openalex updated_date 2026/07/31

Abstract

OBJECTIVE: Machine learning models have utilized semantic features, deep features, or both to assess lung nodule malignancy. However, their reliance on manual annotation during inference, limited interpretability, and sensitivity to imaging variations hinder their application in real-world clinical settings. Thus, this research aims to integrate semantic features derived from radiologists' assessments of nodules, guiding the model to learn clinically relevant, robust, and explainable imaging features for predicting lung cancer. METHODS: We obtained 938 low-dose CT scans from the National Lung Screening Trial (NLST) with 1,261 nodules and semantic features. Additionally, the Lung Image Database Consortium dataset contains 1,018 CT scans, with 2,625 lesions annotated for nodule characteristics. Three external datasets were obtained from UCLA Health, the LUNGx Challenge, and the Duke Lung Cancer Screening. For imaging input, we obtained 2D nodule slices in nine directions from 50×50×50mm nodule crop. We converted structured semantic features into sentences using Gemini. We fine-tuned a pretrained Contrastive Language-Image Pretraining (CLIP) model with a parameter-efficient fine-tuning approach to align imaging and semantic text features and predict the one-year lung cancer diagnosis. RESULTS: Our model outperformed the state-of-the-art (SOTA) models in the NLST test set with an AUROC of 0.901 and AUPRC of 0.776. It also showed robust results in external datasets. Using CLIP, we also obtained predictions on semantic features through zero-shot inference, such as nodule margin (AUROC: 0.807), nodule consistency (0.812), and pleural attachment (0.840). CONCLUSION: By incorporating semantic features into the vision-language model, our approach surpasses the SOTA models in predicting lung cancer from CT scans collected from diverse clinical settings. It provides explainable outputs, aiding clinicians in comprehending the underlying meaning of model predictions. The code is available at https://github.com/luotingzhuang/CLIPnodule.

Citations

Discussions

Related