2020/11/18 by Pablo del Pino, Pablo Pino, Denis Parra +8
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2.7 #I.4.9 #J.3 #Machine Learning (cs.LG) #Radiomics and Machine Learning in Medical Imaging #Topic Modeling #cs.AI #cs.CL #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2011.09257
3 pages, 1 figure, 1 table. Accepted in LatinX in AI workshop at NeurIPS 2020. (v3 updated ack)
openalex publication_date 2020/11/18 · arxiv created 2022/01/15 · arxiv updated 2022/01/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Several deep learning architectures have been proposed over the last years to deal with the problem of generating a written report given an imaging exam as input. Most works evaluate the generated reports using standard Natural Language Processing (NLP) metrics (e.g. BLEU, ROUGE), reporting significant progress. In this article, we contrast this progress by comparing state of the art (SOTA) models against weak baselines. We show that simple and even naive approaches yield near SOTA performance on most traditional NLP metrics. We conclude that evaluation methods in this task should be further studied towards correctly measuring clinical accuracy, ideally involving physicians to contribute to this end.