vix.ing · top · new · best · stats · spec

A Technical Report: Entity Extraction using Both Character-based and Token-based Similarity

2017/02/12 by Zeyi Wen, Dong Deng, Wen, Zeyi +5
Computer Science · Decision Sciences · #Data Quality and Management #Databases (cs.DB) #FOS: Computer and information sciences #Topic Modeling #Web Data Mining and Analysis

paper · pdf · doi:10.48550/arxiv.1702.03519

openalex publication_date 2017/02/12 · openalex created_date 2017/03/16 · openalex updated_date 2026/07/28

Abstract

Entity extraction is fundamental to many text mining tasks such as organisation name recognition. A popular approach to entity extraction is based on matching sub-string candidates in a document against a dictionary of entities. To handle spelling errors and name variations of entities, usually the matching is approximate and edit or Jaccard distance is used to measure dissimilarity between sub-string candidates and the entities. For approximate entity extraction from free text, existing work considers solely character-based or solely token-based similarity and hence cannot simultaneously deal with minor variations at token level and typos. In this paper, we address this problem by considering both character-based similarity and token-based similarity (i.e. two-level similarity). Measuring one-level (e.g. character-based) similarity is computationally expensive, and measuring two-level similarity is dramatically more expensive. By exploiting the properties of the two-level similarity and the weights of tokens, we develop novel techniques to significantly reduce the number of sub-string candidates that require computation of two-level similarity against the dictionary of entities. A comprehensive experimental study on real world datasets show that our algorithm can efficiently extract entities from documents and produce a high F1 score in the range of [0.91, 0.97].

Related