2025/10/14 by Shelley Zixin Shu, Haozhe Luo, Shu, Shelley Zixin +5 · 1 citation
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Boosting (machine learning) #Computer Vision and Pattern Recognition (cs.CV) #Deep learning #FOS: Computer and information sciences #Feature (linguistics) #Feature learning #Generalization #Machine Learning in Healthcare #Multi-task learning #Radiomics and Machine Learning in Medical Imaging #Robustness (evolution) #Seismology and Earthquake Studies #Spurious relationship
paper · pdf · doi:10.48550/arxiv.2510.12704
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/10/14 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
Transformer-based deep learning models have demonstrated exceptional performance in medical imaging by leveraging attention mechanisms for feature representation and interpretability. However, these models are prone to learning spurious correlations, leading to biases and limited generalization. While human-AI attention alignment can mitigate these issues, it often depends on costly manual supervision. In this work, we propose a Hybrid Explanation-Guided Learning (H-EGL) framework that combines self-supervised and human-guided constraints to enhance attention alignment and improve generalization. The self-supervised component of H-EGL leverages class-distinctive attention without relying on restrictive priors, promoting robustness and flexibility. We validate our approach on chest X-ray classification using the Vision Transformer (ViT), where H-EGL outperforms two state-of-the-art Explanation-Guided Learning (EGL) methods, demonstrating superior classification accuracy and generalization capability. Additionally, it produces attention maps that are better aligned with human expertise.