2024/08/29 by Zaiwei Zhang, Zhang, Zaiwei, Gregory P. Meyer +9 · 2 citations
Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Retinal Imaging and Analysis
paper · pdf · doi:10.48550/arxiv.2408.16930
openalex publication_date 2024/08/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
For visual recognition, knowledge distillation typically involves transferring knowledge from a large, well-trained teacher model to a smaller student model. In this paper, we introduce an effective method to distill knowledge from an off-the-shelf vision-language model (VLM), demonstrating that it provides novel supervision in addition to those from a conventional vision-only teacher model. Our key technical contribution is the development of a framework that generates novel text supervision and distills free-form text into a vision encoder. We showcase the effectiveness of our approach, termed VLM-KD, across various benchmark datasets, showing that it surpasses several state-of-the-art long-tail visual classifiers. To our knowledge, this work is the first to utilize knowledge distillation with text supervision generated by an off-the-shelf VLM and apply it to vanilla randomly initialized vision encoders.