2021/11/30 by Carol Anderson, Bo Liu, Anderson, Carol +9
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2111.15641
Submission to the BioCreative VII challenge - Track-3
arxiv created 2021/11/30 · openalex publication_date 2021/11/30 · arxiv updated 2021/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Social media posts contain potentially valuable information about medical conditions and health-related behavior. Biocreative VII Task 3 focuses on mining this information by recognizing mentions of medications and dietary supplements in tweets. We approach this task by fine tuning multiple BERT-style language models to perform token-level classification, and combining them into ensembles to generate final predictions. Our best system consists of five Megatron-BERT-345M models and achieves a strict F1 score of 0.764 on unseen test data.