2020/08/28 by Anh Tuan Nguyen, Nguyen, Anh Tuan
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Misinformation and Its Impacts #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2008.12854
openalex publication_date 2020/08/28 · openalex created_date 2020/09/08 · openalex updated_date 2026/07/28
As the COVID-19 outbreak continues to spread throughout the world, more and more information about the pandemic has been shared publicly on social media. For example, there are a huge number of COVID-19 English Tweets daily on Twitter. However, the majority of those Tweets are uninformative, and hence it is important to be able to automatically select only the informative ones for downstream applications. In this short paper, we present our participation in the W-NUT 2020 Shared Task 2: Identification of Informative COVID-19 English Tweets. Inspired by the recent advances in pretrained Transformer language models, we propose a simple yet effective baseline for the task. Despite its simplicity, our proposed approach shows very competitive results in the leaderboard as we ranked 8 over 56 teams participated in total.