vix.ing · top · new · best · stats · spec

Toward Network-based Keyword Extraction from Multitopic Web Documents

2014/07/14 by Sabina Šišović, Šišović, Sabina, Sanda Martinčić-Ipšić +3
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR)

paper · pdf · doi:10.48550/arxiv.1407.3636

openalex publication_date 2014/07/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper we analyse the selectivity measure calculated from the complex network in the task of the automatic keyword extraction. Texts, collected from different web sources (portals, forums), are represented as directed and weighted co-occurrence complex networks of words. Words are nodes and links are established between two nodes if they are directly co-occurring within the sentence. We test different centrality measures for ranking nodes - keyword candidates. The promising results are achieved using the selectivity measure. Then we propose an approach which enables extracting word pairs according to the values of the in/out selectivity and weight measures combined with filtering.

Related