2021/02/13 by Julia Wu, Venkatesh Sivaraman, Wu, Julia +7
Computer Science · Social Sciences · #Computers and Society (cs.CY) #FOS: Computer and information sciences #Misinformation and Its Impacts #Social Media in Health Education #Social and Information Networks (cs.SI) #Topic Modeling #cs.CY #cs.SI
paper · pdf · doi:10.48550/arxiv.2102.06836
24 pages, 5 figures. To be published in the Journal of Biomedical Informatics
openalex publication_date 2021/02/13 · arxiv created 2021/06/28 · arxiv updated 2021/06/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The rapid evolution of the COVID-19 pandemic has underscored the need to quickly disseminate the latest clinical knowledge during a public-health emergency. One surprisingly effective platform for healthcare professionals (HCPs) to share knowledge and experiences from the front lines has been social media (for example, the "#medtwitter" community on Twitter). However, identifying clinically-relevant content in social media without manual labeling is a challenge because of the sheer volume of irrelevant data. We present an unsupervised, iterative approach to mine clinically relevant information from social media data, which begins by heuristically filtering for HCP-authored texts and incorporates topic modeling and concept extraction with MetaMap. This approach identifies granular topics and tweets with high clinical relevance from a set of about 52 million COVID-19-related tweets from January to mid-June 2020. We also show that because the technique does not require manual labeling, it can be used to identify emerging topics on a week-to-week basis. Our method can aid in future public-health emergencies by facilitating knowledge transfer among healthcare workers in a rapidly-changing information environment, and by providing an efficient and unsupervised way of highlighting potential areas for clinical research.