vix.ing · top · new · best · stats · spec

COVID-Twitter-BERT: A Natural Language Processing Model to Analyse COVID-19 Content on Twitter

2020/05/15 by Martín Müller, Müller, Martin, Marcel Salathé +3 · 9 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Misinformation and Its Impacts #Sentiment Analysis and Opinion Mining #Social and Information Networks (cs.SI)

paper · pdf · doi:10.48550/arxiv.2005.07503

openalex publication_date 2020/05/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this work, we release COVID-Twitter-BERT (CT-BERT), a transformer-based model, pretrained on a large corpus of Twitter messages on the topic of COVID-19. Our model shows a 10-30% marginal improvement compared to its base model, BERT-Large, on five different classification datasets. The largest improvements are on the target domain. Pretrained transformer models, such as CT-BERT, are trained on a specific target domain and can be used for a wide variety of natural language processing tasks, including classification, question-answering and chatbots. CT-BERT is optimised to be used on COVID-19 content, in particular social media posts from Twitter.

Cited by

Related