vix.ing · top · new · best · stats · spec

Balancing the composition of word embeddings across heterogenous data\n sets

2020/01/14 by Stephanie Brandl, Brandl, Stephanie, David Lassner +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Text and Document Classification Technologies #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2001.04693

openalex publication_date 2020/01/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Word embeddings capture semantic relationships based on contextual\ninformation and are the basis for a wide variety of natural language processing\napplications. Notably these relationships are solely learned from the data and\nsubsequently the data composition impacts the semantic of embeddings which\narguably can lead to biased word vectors. Given qualitatively different data\nsubsets, we aim to align the influence of single subsets on the resulting word\nvectors, while retaining their quality. In this regard we propose a criteria to\nmeasure the shift towards a single data subset and develop approaches to meet\nboth objectives. We find that a weighted average of the two subset embeddings\nbalances the influence of those subsets while word similarity performance\ndecreases. We further propose a promising optimization approach to balance\ninfluences and quality of word embeddings.\n

Citations

Related