2017/09/06 by Shashank Gupta, Sachin Pawar, Gupta, Shashank +7
Biochemistry, Genetics and Molecular Biology · Computer Science · Social Sciences · #Academic integrity and plagiarism #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Computational Drug Discovery Methods #FOS: Computer and information sciences #Information Retrieval (cs.IR)
paper · pdf · doi:10.48550/arxiv.1709.01687
openalex publication_date 2017/09/06 · openalex created_date 2022/10/07 · openalex updated_date 2026/07/28
Social media is an useful platform to share health-related information due to\nits vast reach. This makes it a good candidate for public-health monitoring\ntasks, specifically for pharmacovigilance. We study the problem of extraction\nof Adverse-Drug-Reaction (ADR) mentions from social media, particularly from\ntwitter. Medical information extraction from social media is challenging,\nmainly due to short and highly information nature of text, as compared to more\ntechnical and formal medical reports.\n Current methods in ADR mention extraction relies on supervised learning\nmethods, which suffers from labeled data scarcity problem. The State-of-the-art\nmethod uses deep neural networks, specifically a class of Recurrent Neural\nNetwork (RNN) which are Long-Short-Term-Memory networks (LSTMs)\n citehochreiter1997long. Deep neural networks, due to their large number of\nfree parameters relies heavily on large annotated corpora for learning the end\ntask. But in real-world, it is hard to get large labeled data, mainly due to\nheavy cost associated with manual annotation. Towards this end, we propose a\nnovel semi-supervised learning based RNN model, which can leverage unlabeled\ndata also present in abundance on social media. Through experiments we\ndemonstrate the effectiveness of our method, achieving state-of-the-art\nperformance in ADR mention extraction.\n