2024/07/05 by Nikolaos Giarelis, Nikos Karacapilidis · 15 citations
Computer Science · #Advanced Text Analysis Techniques #Artificial intelligence #Categorization #Computer science #Context (archaeology) #Embedding #Field (mathematics) #Focus (optics) #Information retrieval #Natural language processing #Salient #Sentence #Word embedding
paper · pdf · doi:10.1007/s10115-024-02164-w
published in Knowledge and Information Systems 66(11), 6493-6526 (Springer Science+Business Media)
crossref issued 2024/07/05 · crossref published 2024/07/05 · crossref published-online 2024/07/05 · openalex publication_date 2024/07/05 · crossref created 2024/07/05 · crossref deposited 2024/09/28 · crossref published-print 2024/11/01 · openalex created_date 2025/10/10 · crossref indexed 2026/08/04 · openalex updated_date 2026/08/05
Abstract Keyphrase extraction is a subtask of natural language processing referring to the automatic extraction of salient terms that semantically capture the key themes and topics of a document. Earlier literature reviews focus on classical approaches that employ various statistical or graph-based techniques; these approaches miss important keywords/keyphrases, due to their inability to fully utilize context (that is present or not) in a document, thus achieving low F1 scores. Recent advances in deep learning and word/sentence embedding vectors lead to the development of new approaches, which address the lack of context and outperform the majority of classical ones. Taking the above into account, the contribution of this review is fourfold: (i) we analyze the state-of-the-art keyphrase extraction approaches and categorize them upon their employed techniques; (ii) we provide a comparative evaluation of these approaches, using well-known datasets of the literature and popular evaluation metrics, such as the F1 score; (iii) we provide a series of insights on various keyphrase extraction issues, including alternative approaches and future research directions; (iv) we make the datasets and code used in our experiments public, aiming to further increase the reproducibility of this work and facilitate future research in the field.