2020/11/06 by Mohit Bhardwaj, Bhardwaj, Mohit, Md Shad Akhtar +7 · 7 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Spam and Phishing Detection
paper · pdf · doi:10.48550/arxiv.2011.03588
openalex publication_date 2020/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we present a novel hostility detection dataset in Hindi language. We collect and manually annotate ~8200 online posts. The annotated dataset covers four hostility dimensions: fake news, hate speech, offensive, and defamation posts, along with a non-hostile label. The hostile posts are also considered for multi-label tags due to a significant overlap among the hostile classes. We release this dataset as part of the CONSTRAINT-2021 shared task on hostile post detection.