2021/01/12 by Zhi Hong, J. Gregory Pauloski, Hong, Zhi +9
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational Drug Discovery Methods #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Misinformation and Its Impacts #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2101.04617
openalex publication_date 2021/01/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Researchers worldwide are seeking to repurpose existing drugs or discover new drugs to counter the disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). A promising source of candidates for such studies is molecules that have been reported in the scientific literature to be drug-like in the context of coronavirus research. We report here on a project that leverages both human and artificial intelligence to detect references to drug-like molecules in free text. We engage non-expert humans to create a corpus of labeled text, use this labeled corpus to train a named entity recognition model, and employ the trained model to extract 10912 drug-like molecules from the COVID-19 Open Research Dataset Challenge (CORD-19) corpus of 198875 papers. Performance analyses show that our automated extraction model can achieve performance on par with that of non-expert humans.