vix.ing · top · new · best · stats · spec

Do Language Embeddings Capture Scales?

2020/10/11 by Xikun Zhang, Deepak Ramachandran, Zhang, Xikun +7
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2010.05345

openalex publication_date 2020/10/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Pretrained Language Models (LMs) have been shown to possess significant linguistic, common sense, and factual knowledge. One form of knowledge that has not been studied yet in this context is information about the scalar magnitudes of objects. We show that pretrained language models capture a significant amount of this information but are short of the capability required for general common-sense reasoning. We identify contextual information in pre-training and numeracy as two key factors affecting their performance and show that a simple method of canonicalizing numbers can have a significant effect on the results.

Citations

Related