2021/11/26 by Dongfang Xu, Xu, Dongfang, Shan Chen +4
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Artificial intelligence #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Computer science #Engineering #Ensemble learning #F1 score #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine learning #Named-entity recognition #Natural Language Processing Techniques #Natural language processing #Ranking (information retrieval) #Task (project management) #Topic Modeling #Track (disk drive) #Transformer #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2111.13726
Proceedings of the seventh BioCreative challenge evaluation workshop (https://biocreative.bioinformatics.udel.edu/resources/publications/bc-vii-workshop-proceedings/)
arxiv created 2021/11/26 · openalex publication_date 2021/11/26 · arxiv updated 2021/11/30 · openalex created_date 2021/12/06 · openalex updated_date 2026/08/06
In this paper, we present our work participating in the BioCreative VII Track 3 - automatic extraction of medication names in tweets, where we implemented a multi-task learning model that is jointly trained on text classification and sequence labelling. Our best system run achieved a strict F1 of 80.4, ranking first and more than 10 points higher than the average score of all participants. Our analyses show that the ensemble technique, multi-task learning, and data augmentation are all beneficial for medication detection in tweets.