2014/01/25 by Shibamouli Lahiri, Sagnik Ray Choudhury, Lahiri, Shibamouli +3 · 1 citation
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Information Retrieval and Search Behavior
paper · pdf · doi:10.48550/arxiv.1401.6571
openalex publication_date 2014/01/25 · openalex created_date 2025/10/24 · openalex updated_date 2026/07/28
Keyword and keyphrase extraction is an important problem in natural language\nprocessing, with applications ranging from summarization to semantic search to\ndocument clustering. Graph-based approaches to keyword and keyphrase extraction\navoid the problem of acquiring a large in-domain training corpus by applying\nvariants of PageRank algorithm on a network of words. Although graph-based\napproaches are knowledge-lean and easily adoptable in online systems, it\nremains largely open whether they can benefit from centrality measures other\nthan PageRank. In this paper, we experiment with an array of centrality\nmeasures on word and noun phrase collocation networks, and analyze their\nperformance on four benchmark datasets. Not only are there centrality measures\nthat perform as well as or better than PageRank, but they are much simpler\n(e.g., degree, strength, and neighborhood size). Furthermore, centrality-based\nmethods give results that are competitive with and, in some cases, better than\ntwo strong unsupervised baselines.\n