vix.ing · top · new · best · stats · spec

Multi-Document Keyphrase Extraction: Dataset, Baselines and Review

2021/10/03 by Ori Shapira, Ramakanth Pasunuru, Shapira, Ori +5
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2110.01073

openalex publication_date 2021/10/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Keyphrase extraction has been extensively researched within the single-document setting, with an abundance of methods, datasets and applications. In contrast, multi-document keyphrase extraction has been infrequently studied, despite its utility for describing sets of documents, and its use in summarization. Moreover, no prior dataset exists for multi-document keyphrase extraction, hindering the progress of the task. Recent advances in multi-text processing make the task an even more appealing challenge to pursue. To stimulate this pursuit, we present here the first dataset for the task, MK-DUC-01, which can serve as a new benchmark, and test multiple keyphrase extraction baselines on our data. In addition, we provide a brief, yet comprehensive, literature review of the task.

Related