2016/09/02 by Bridget T. McInnes, McInnes, Bridget T., Ted Pedersen +1
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1609.00559
openalex publication_date 2016/09/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Vector space methods that measure semantic similarity and relatedness often\nrely on distributional information such as co--occurrence frequencies or\nstatistical measures of association to weight the importance of particular\nco--occurrences. In this paper, we extend these methods by incorporating a\nmeasure of semantic similarity based on a human curated taxonomy into a\nsecond--order vector representation. This results in a measure of semantic\nrelatedness that combines both the contextual information available in a\ncorpus--based vector space representation with the semantic knowledge found in\na biomedical ontology. Our results show that incorporating semantic similarity\ninto a second order co--occurrence matrices improves correlation with human\njudgments for both similarity and relatedness, and that our method compares\nfavorably to various different word embedding methods that have recently been\nevaluated on the same reference standards we have used.\n