vix.ing · top · new · best · stats · spec

Measuring research interest similarity with transition probabilities

2025/01/01 by Attila Varga, Sadamori Kojaku, Filipi N. Silva · 1 voice
Decision Sciences · Biochemistry, Genetics and Molecular Biology · #Scientific Computing and Data Management #scientometrics and bibliometrics research #Biomedical Text Mining and Ontologies

paper · pdf · doi:10.1162/qss.a.13

openalex publication_date 2025/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/26

Abstract

Abstract We introduce a family of paper and author similarity measures based on the concept that papers are more similar if they are more likely to be retrieved during a literature search following backward and forward citations. As this browsing process resembles a walk in a citation network, we operationalize the concept using the transition probability (TP) of random walkers. The proposed measures are continuous and symmetric, and can be implemented on any citation network. We conduct validation tests of the TP concept and other extant alternatives to gauge which metric can classify papers and predict future coauthors most consistently across different scales of analysis (coauthorships, journals, and disciplines). Our results show that the proposed basic TP measure outperforms alternative metrics such as personalized PageRank and the node2vec machine-learning technique in classification tasks at various scales. Additionally, we discuss how publication-level data can be leveraged to approximate the research interest similarity of individual scientists. This paper is accompanied by a Python package that implements all the tested metrics.

Citations

Discussions

Related