vix.ing · top · new · best · stats · spec

ArCOV19-Rumors: Arabic COVID-19 Twitter Dataset for Misinformation Detection

2020/10/17 by Fatima Haouari, Haouari, Fatima, Maram Hasanain +5 · 1 citation
Computer Science · Social Sciences · #Arabic #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Computer security #Coronavirus disease 2019 (COVID-19) #Disinformation #Entertainment #Exploit #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Information Retrieval (cs.IR) #Information retrieval #Internet privacy #Linguistics #Media Influence and Politics #Misinformation #Misinformation and Its Impacts #Natural language processing #Political science #Sentiment analysis #Social and Information Networks (cs.SI) #Social media #Spam and Phishing Detection #World Wide Web #cs.CL #cs.IR #cs.SI

paper · pdf · doi:10.48550/arxiv.2010.08768

This work was accepted at the Sixth Arabic Natural Language Processing Workshop (EACL/WANLP 2021)

openalex publication_date 2020/10/17 · arxiv created 2021/03/13 · arxiv updated 2021/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

In this paper we introduce ArCOV19-Rumors, an Arabic COVID-19 Twitter dataset for misinformation detection composed of tweets containing claims from 27th January till the end of April 2020. We collected 138 verified claims, mostly from popular fact-checking websites, and identified 9.4K relevant tweets to those claims. Tweets were manually-annotated by veracity to support research on misinformation detection, which is one of the major problems faced during a pandemic. ArCOV19-Rumors supports two levels of misinformation detection over Twitter: verifying free-text claims (called claim-level verification) and verifying claims expressed in tweets (called tweet-level verification). Our dataset covers, in addition to health, claims related to other topical categories that were influenced by COVID-19, namely, social, politics, sports, entertainment, and religious. Moreover, we present benchmarking results for tweet-level verification on the dataset. We experimented with SOTA models of versatile approaches that either exploit content, user profiles features, temporal features and propagation structure of the conversational threads for tweet verification.

Cited by

Related