2021/04/29 by Aina Garí Soler, Soler, Aina Garí, Marianna Apidianaki +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2104.14694
openalex publication_date 2021/04/29 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Pre-trained language models (LMs) encode rich information about linguistic\nstructure but their knowledge about lexical polysemy remains unclear. We\npropose a novel experimental setup for analysing this knowledge in LMs\nspecifically trained for different languages (English, French, Spanish and\nGreek) and in multilingual BERT. We perform our analysis on datasets carefully\ndesigned to reflect different sense distributions, and control for parameters\nthat are highly correlated with polysemy such as frequency and grammatical\ncategory. We demonstrate that BERT-derived representations reflect words'\npolysemy level and their partitionability into senses. Polysemy-related\ninformation is more clearly present in English BERT embeddings, but models in\nother languages also manage to establish relevant distinctions between words at\ndifferent polysemy levels. Our results contribute to a better understanding of\nthe knowledge encoded in contextualised representations and open up new avenues\nfor multilingual lexical semantics research.\n