vix.ing · top · new · best · stats · spec

Improving Correlation with Human Judgments by Integrating Semantic\n Similarity with Second--Order Vectors

2016/09/02 by Bridget T. McInnes, McInnes, Bridget T., Ted Pedersen +1
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1609.00559

openalex publication_date 2016/09/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Vector space methods that measure semantic similarity and relatedness often\nrely on distributional information such as co--occurrence frequencies or\nstatistical measures of association to weight the importance of particular\nco--occurrences. In this paper, we extend these methods by incorporating a\nmeasure of semantic similarity based on a human curated taxonomy into a\nsecond--order vector representation. This results in a measure of semantic\nrelatedness that combines both the contextual information available in a\ncorpus--based vector space representation with the semantic knowledge found in\na biomedical ontology. Our results show that incorporating semantic similarity\ninto a second order co--occurrence matrices improves correlation with human\njudgments for both similarity and relatedness, and that our method compares\nfavorably to various different word embedding methods that have recently been\nevaluated on the same reference standards we have used.\n

Related