2020/06/08 by Frances Adriana Laureano De Leon, De Leon, Frances Adriana Laureano, Florimond Guéniat +3
Computer Science · #Computation and Language (cs.CL) #Digital Communication and Language #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Text Readability and Simplification
paper · pdf · doi:10.48550/arxiv.2006.04597
openalex publication_date 2020/06/08 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
The growing popularity and applications of sentiment analysis of social media\nposts has naturally led to sentiment analysis of posts written in multiple\nlanguages, a practice known as code-switching. While recent research into\ncode-switched posts has focused on the use of multilingual word embeddings,\nthese embeddings were not trained on code-switched data. In this work, we\npresent word-embeddings trained on code-switched tweets, specifically those\nthat make use of Spanish and English, known as Spanglish. We explore the\nembedding space to discover how they capture the meanings of words in both\nlanguages. We test the effectiveness of these embeddings by participating in\nSemEval 2020 Task 9: ~\Sentiment Analysis on Code-Mixed Social Media\nText. We utilised them to train a sentiment classifier that achieves an F-1\nscore of 0.722. This is higher than the baseline for the competition of 0.656,\nwith our team (codalab username \francesita) ranking 14 out of 29\nparticipating teams, beating the baseline.\n