vix.ing · top · new · best · stats

Misinformation detection in Luganda-English code-mixed social media text

2021/03/31 by Peter Nabende, Nabende, Peter, David Kabiito +8 · 2 citations
Computer Science · Social Sciences · #Artificial intelligence #Code (set theory) #Computation and Language (cs.CL) #Computer science #Computer security #Discriminative model #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Information retrieval #Machine learning #Misinformation #Misinformation and Its Impacts #Natural language processing #Set (abstract data type) #Social and Information Networks (cs.SI) #Social media #Spam and Phishing Detection #Support vector machine #World Wide Web #cs.CL #cs.SI

paper · pdf · doi:10.48550/arxiv.2104.00124

published in arXiv (Cornell University) (Cornell University) · Accepted at African NLP workshop @EACL 2021

openalex publication_date 2021/03/31 · arxiv created 2021/04/03 · arxiv updated 2021/04/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

The increasing occurrence, forms, and negative effects of misinformation on social media platforms has necessitated more misinformation detection tools. Currently, work is being done addressing COVID-19 misinformation however, there are no misinformation detection tools for any of the 40 distinct indigenous Ugandan languages. This paper addresses this gap by presenting basic language resources and a misinformation detection data set based on code-mixed Luganda-English messages sourced from the Facebook and Twitter social media platforms. Several machine learning methods are applied on the misinformation detection data set to develop classification models for detecting whether a code-mixed Luganda-English message contains misinformation or not. A 10-fold cross validation evaluation of the classification methods in an experimental misinformation detection task shows that a Discriminative Multinomial Naive Bayes (DMNB) method achieves the highest accuracy and F-measure of 78.19% and 77.90% respectively. Also, Support Vector Machine and Bagging ensemble classification models achieve comparable results. These results are promising since the machine learning models are based on n-gram features from only the misinformation detection dataset.

Citations

Related