2021/06/28 by Florian Boudin, Béatrice Daille, Boudin, Florian +5
Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR)
paper · pdf · doi:10.48550/arxiv.2106.14731
openalex publication_date 2021/06/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Scientific digital libraries play a critical role in the development and dissemination of scientific literature. Despite dedicated search engines, retrieving relevant publications from the ever-growing body of scientific literature remains challenging and time-consuming. Indexing scientific articles is indeed a difficult matter, and current models solely rely on a small portion of the articles (title and abstract) and on author-assigned keyphrases when available. This results in a frustratingly limited access to scientific knowledge. The goal of the DELICES project is to address this pitfall by exploiting semantic relations between scientific articles to both improve and enrich indexing. To this end, we will rely on the latest advances in semantic representations to both increase the relevance of keyphrases extracted from the documents, and extend indexing to new terms borrowed from semantically similar documents.