2021/12/07 by Naman Paharia, Paharia, Naman, Muhammad Syafiq Mohd Pozi +3
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Semantic Web and Ontologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2112.03634
openalex publication_date 2021/12/07 · openalex created_date 2022/11/13 · openalex updated_date 2026/07/28
The amount of scholarly data has been increasing dramatically over the last\nyears. For newcomers to a particular science domain (e.g., IR, physics, NLP) it\nis often difficult to spot larger trends and to position the latest research in\nthe context of prior scientific achievements and breakthroughs. Similarly,\nresearchers in the history of science are interested in tools that allow them\nto analyze and visualize changes in particular scientific domains. Temporal\nsummarization and related methods should be then useful for making sense of\nlarge volumes of scientific discourse data aggregated over time. We demonstrate\na novel approach to analyze the collections of research papers published over\nlonger time periods to provide a high-level overview of important semantic\nchanges that occurred over the progress of time. Our approach is based on\ncomparing word semantic representations over time and aims to support users in\na better understanding of large domain-focused archives of scholarly\npublications. As an example dataset we use the ACL Anthology Reference Corpus\nthat spans from 1979 to 2015 and contains 22,878 scholarly articles.\n