vix.ing · top · new · best · stats · spec

Closing the Gap: Joint De-Identification and Concept Extraction in the\n Clinical Domain

2020/05/19 by Lukas Lange, Lange, Lukas, Heike Adel +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2005.09397

openalex publication_date 2020/05/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Exploiting natural language processing in the clinical domain requires\nde-identification, i.e., anonymization of personal information in texts.\nHowever, current research considers de-identification and downstream tasks,\nsuch as concept extraction, only in isolation and does not study the effects of\nde-identification on other tasks. In this paper, we close this gap by reporting\nconcept extraction performance on automatically anonymized data and\ninvestigating joint models for de-identification and concept extraction. In\nparticular, we propose a stacked model with restricted access to\nprivacy-sensitive information and a multitask model. We set the new state of\nthe art on benchmark datasets in English (96.1% F1 for de-identification and\n88.9% F1 for concept extraction) and Spanish (91.4% F1 for concept extraction).\n

Cited by

Related