vix.ing · top · new · best · stats · spec

A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports

2025/04/28 by Henning Schäfer, Schäfer, Henning, Cynthia Sabrina Schmidt +17
Computer Science · Health Professions · #68T07 #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Electronic Health Records Systems #FOS: Computer and information sciences #H.3.3 #Handwritten Text Recognition Techniques #I.2.7 #I.4.7 #I.7.5 #J.3 #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2504.20220

openalex publication_date 2025/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Despite the growing adoption of electronic health records, many processes still rely on paper documents, reflecting the heterogeneous real-world conditions in which healthcare is delivered. The manual transcription process is time-consuming and prone to errors when transferring paper-based data to digital formats. To streamline this workflow, this study presents an open-source pipeline that extracts and categorizes checkbox data from scanned documents. Demonstrated on transfusion reaction reports, the design supports adaptation to other checkbox-rich document types. The proposed method integrates checkbox detection, multilingual optical character recognition (OCR) and multilingual vision-language models (VLMs). The pipeline achieves high precision and recall compared against annually compiled gold-standards from 2017 to 2024. The result is a reduction in administrative workload and accurate regulatory reporting. The open-source availability of this pipeline encourages self-hosted parsing of checkbox forms.

Citations

Related