2021/02/24 by Akshat Gupta, Gupta, Akshat, Sai Krishna Rallabandi +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2102.12407
openalex publication_date 2021/02/24 · openalex created_date 2021/03/01 · openalex updated_date 2026/07/28
Using task-specific pre-training and leveraging cross-lingual transfer are two of the most popular ways to handle code-switched data. In this paper, we aim to compare the effects of both for the task of sentiment analysis. We work with two Dravidian Code-Switched languages - Tamil-Engish and Malayalam-English and four different BERT based models. We compare the effects of task-specific pre-training and cross-lingual transfer and find that task-specific pre-training results in superior zero-shot and supervised performance when compared to performance achieved by leveraging cross-lingual transfer from multilingual BERT models.