vix.ing · top · new · best · stats

Closing the Gap: Joint De-Identification and Concept Extraction in the Clinical Domain

2020/05/19 by Lukas Lange, Lange, Lukas, Heike Adel +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2005.09397

ACL 2020

arxiv created 2020/05/19 · openalex publication_date 2020/05/19 · arxiv updated 2020/05/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Exploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept extraction, only in isolation and does not study the effects of de-identification on other tasks. In this paper, we close this gap by reporting concept extraction performance on automatically anonymized data and investigating joint models for de-identification and concept extraction. In particular, we propose a stacked model with restricted access to privacy-sensitive information and a multitask model. We set the new state of the art on benchmark datasets in English (96.1% F1 for de-identification and 88.9% F1 for concept extraction) and Spanish (91.4% F1 for concept extraction).

Cited by

Related