vix.ing · top · new · best · stats · spec

XLEnt: Mining a Large Cross-lingual Entity Dataset with\n Lexical-Semantic-Phonetic Word Alignment

2021/04/17 by Ahmed El-Kishky, Adithya Renduchintala, El-Kishky, Ahmed +8 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2104.08597

openalex publication_date 2021/04/17 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Cross-lingual named-entity lexica are an important resource to multilingual\nNLP tasks such as machine translation and cross-lingual wikification. While\nknowledge bases contain a large number of entities in high-resource languages\nsuch as English and French, corresponding entities for lower-resource languages\nare often missing. To address this, we propose Lexical-Semantic-Phonetic Align\n(LSP-Align), a technique to automatically mine cross-lingual entity lexica from\nmined web data. We demonstrate LSP-Align outperforms baselines at extracting\ncross-lingual entity pairs and mine 164 million entity pairs from 120 different\nlanguages aligned with English. We release these cross-lingual entity pairs\nalong with the massively multilingual tagged named entity corpus as a resource\nto the NLP community.\n

Cited by

Related