vix.ing · top · new · best · stats · spec

GGPONC: A Corpus of German Medical Text with Rich Metadata Based on\n Clinical Practice Guidelines

2020/07/13 by Florian Borchert, Christina Lohr, Borchert, Florian +13
Arts and Humanities · Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #linguistics and terminology studies

paper · pdf · doi:10.48550/arxiv.2007.06400

openalex publication_date 2020/07/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The lack of publicly accessible text corpora is a major obstacle for progress\nin natural language processing. For medical applications, unfortunately, all\nlanguage communities other than English are low-resourced. In this work, we\npresent GGPONC (German Guideline Program in Oncology NLP Corpus), a freely\ndistributable German language corpus based on clinical practice guidelines for\noncology. This corpus is one of the largest ever built from German medical\ndocuments. Unlike clinical documents, clinical guidelines do not contain any\npatient-related information and can therefore be used without data protection\nrestrictions. Moreover, GGPONC is the first corpus for the German language\ncovering diverse conditions in a large medical subfield and provides a variety\nof metadata, such as literature references and evidence levels. By applying and\nevaluating existing medical information extraction pipelines for German text,\nwe are able to draw comparisons for the use of medical language to other\ncorpora, medical and non-medical ones.\n

Related