vix.ing · top · new · best · stats · spec

Walk in Wild: An Ensemble Approach for Hostility Detection in Hindi Posts

2021/01/15 by Chander Shekhar, Shekhar, Chander, Bhavya Bagla +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Sentiment Analysis and Opinion Mining #Spam and Phishing Detection

paper · pdf · doi:10.48550/arxiv.2101.06004

openalex publication_date 2021/01/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

As the reach of the internet increases, pejorative terms started flooding over social media platforms. This leads to the necessity of identifying hostile content on social media platforms. Identification of hostile contents on low-resource languages like Hindi poses different challenges due to its diverse syntactic structure compared to English. In this paper, we develop a simple ensemble based model on pre-trained mBERT and popular classification algorithms like Artificial Neural Network (ANN) and XGBoost for hostility detection in Hindi posts. We formulated this problem as binary classification (hostile and non-hostile class) and multi-label multi-class classification problem (for more fine-grained hostile classes). We received third overall rank in the competition and weighted F1-scores of ~0.969 and ~0.61 on the binary and multi-label multi-class classification tasks respectively.

Citations

Related