vix.ing · top · new · best · stats · spec

Bad Form: Comparing Context-Based and Form-Based Few-Shot Learning in\n Distributional Semantic Models

2019/10/01 by Jeroen Van Hautte, Van Hautte, Jeroen, Guy Emerson +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1910.00275

openalex publication_date 2019/10/01 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28

Abstract

Word embeddings are an essential component in a wide range of natural\nlanguage processing applications. However, distributional semantic models are\nknown to struggle when only a small number of context sentences are available.\nSeveral methods have been proposed to obtain higher-quality vectors for these\nwords, leveraging both this context information and sometimes the word forms\nthemselves through a hybrid approach. We show that the current tasks do not\nsuffice to evaluate models that use word-form information, as such models can\neasily leverage word forms in the training data that are related to word forms\nin the test data. We introduce 3 new tasks, allowing for a more balanced\ncomparison between models. Furthermore, we show that hyperparameters that have\nlargely been ignored in previous work can consistently improve the performance\nof both baseline and advanced models, achieving a new state of the art on 4 out\nof 6 tasks.\n

Related