vix.ing · top · new · best · stats · spec

Improving Sentiment Analysis over non-English Tweets using Multilingual\n Transformers and Automatic Translation for Data-Augmentation

2020/10/07 by Valentin Barrière, Barriere, Valentin, Alexandra Balahur +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2010.03486

openalex publication_date 2020/10/07 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Tweets are specific text data when compared to general text. Although\nsentiment analysis over tweets has become very popular in the last decade for\nEnglish, it is still difficult to find huge annotated corpora for non-English\nlanguages. The recent rise of the transformer models in Natural Language\nProcessing allows to achieve unparalleled performances in many tasks, but these\nmodels need a consequent quantity of text to adapt to the tweet domain. We\npropose the use of a multilingual transformer model, that we pre-train over\nEnglish tweets and apply data-augmentation using automatic translation to adapt\nthe model to non-English languages. Our experiments in French, Spanish, German\nand Italian suggest that the proposed technique is an efficient way to improve\nthe results of the transformers over small corpora of tweets in a non-English\nlanguage.\n

Cited by

Related