vix.ing · top · new · best · stats · spec

Hybrid Approaches for our Participation to the n2c2 Challenge on Cohort\n Selection for Clinical Trials

2019/03/19 by Xavier Tannier, Tannier, Xavier, Nicolás Paris +24
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1903.07879

openalex publication_date 2019/03/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Objective: Natural language processing can help minimize human intervention\nin identifying patients meeting eligibility criteria for clinical trials, but\nthere is still a long way to go to obtain a general and systematic approach\nthat is useful for researchers. We describe two methods taking a step in this\ndirection and present their results obtained during the n2c2 challenge on\ncohort selection for clinical trials. Materials and Methods: The first method\nis a weakly supervised method using an unlabeled corpus (MIMIC) to build a\nsilver standard, by producing semi-automatically a small and very precise set\nof rules to detect some samples of positive and negative patients. This silver\nstandard is then used to train a traditional supervised model. The second\nmethod is a terminology-based approach where a medical expert selects the\nappropriate concepts, and a procedure is defined to search the terms and check\nthe structural or temporal constraints. Results: On the n2c2 dataset containing\nannotated data about 13 selection criteria on 288 patients, we obtained an\noverall F1-measure of 0.8969, which is the third best result out of 45\nparticipant teams, with no statistically significant difference with the\nbest-ranked team. Discussion: Both approaches obtained very encouraging results\nand apply to different types of criteria. The weakly supervised method requires\nexplicit descriptions of positive and negative examples in some reports. The\nterminology-based method is very efficient when medical concepts carry most of\nthe relevant information. Conclusion: It is unlikely that much more annotated\ndata will be soon available for the task of identifying a wide range of patient\nphenotypes. One must focus on weakly or non-supervised learning methods using\nboth structured and unstructured data and relying on a comprehensive\nrepresentation of the patients.\n

Related