vix.ing · top · new · best · stats · spec

Multilingual Abusiveness Identification on Code-Mixed Social Media Text

2022/03/01 by Ekagra Ranjan, Ranjan, Ekagra, Naman Poddar +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social and Information Networks (cs.SI) #Text Readability and Simplification

paper · pdf · doi:10.48550/arxiv.2204.01848

openalex publication_date 2022/03/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Social Media platforms have been seeing adoption and growth in their usage over time. This growth has been further accelerated with the lockdown in the past year when people's interaction, conversation, and expression were limited physically. It is becoming increasingly important to keep the platform safe from abusive content for better user experience. Much work has been done on English social media content but text analysis on non-English social media is relatively underexplored. Non-English social media content have the additional challenges of code-mixing, transliteration and using different scripture in same sentence. In this work, we propose an approach for abusiveness identification on the multilingual Moj dataset which comprises of Indic languages. Our approach tackles the common challenges of non-English social media content and can be extended to other languages as well.

Related