vix.ing · top · new · best · stats · spec

Building Multilingual Corpora for a Complex Named Entity Recognition and\n Classification Hierarchy using Wikipedia and DBpedia

2022/12/14 by Diego Alves, Gaurish Thakkar, Alves, Diego +7
Computer Science · Social Sciences · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Wikis in Education and Collaboration

paper · pdf · doi:10.48550/arxiv.2212.07429

openalex publication_date 2022/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

With the ever-growing popularity of the field of NLP, the demand for datasets\nin low resourced-languages follows suit. Following a previously established\nframework, in this paper, we present the UNER dataset, a multilingual and\nhierarchical parallel corpus annotated for named-entities. We describe in\ndetail the developed procedure necessary to create this type of dataset in any\nlanguage available on Wikipedia with DBpedia information. The three-step\nprocedure extracts entities from Wikipedia articles, links them to DBpedia, and\nmaps the DBpedia sets of classes to the UNER labels. This is followed by a\npost-processing procedure that significantly increases the number of identified\nentities in the final results. The paper concludes with a statistical and\nqualitative analysis of the resulting dataset.\n

Related