2018/10/25 by Sven Buechel, João Sedoc, Buechel, Sven +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1810.10949
Published at PEOPLES 2020
openalex publication_date 2018/10/25 · arxiv created 2020/12/07 · arxiv updated 2020/12/08 · openalex created_date 2022/08/02 · openalex updated_date 2026/07/28
One of the major downsides of Deep Learning is its supposed need for vast amounts of training data. As such, these techniques appear ill-suited for NLP areas where annotated data is limited, such as less-resourced languages or emotion analysis, with its many nuanced and hard-to-acquire annotation formats. We conduct a questionnaire study indicating that indeed the vast majority of researchers in emotion analysis deems neural models inferior to traditional machine learning when training data is limited. In stark contrast to those survey results, we provide empirical evidence for English, Polish, and Portuguese that commonly used neural architectures can be trained on surprisingly few observations, outperforming n-gram based ridge regression on only 100 data points. Our analysis suggests that high-quality, pre-trained word embeddings are a main factor for achieving those results.