2019/11/21 by Zeyi Wen, Zeyu Huang, Wen, Zeyi +3
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling #Web Data Mining and Analysis
paper · pdf · doi:10.48550/arxiv.1911.09373
openalex publication_date 2019/11/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Entity extraction is an important task in text mining and natural language processing. A popular method for entity extraction is by comparing substrings from free text against a dictionary of entities. In this paper, we present several techniques as a post-processing step for improving the effectiveness of the existing entity extraction technique. These techniques utilise models trained with the web-scale corpora which makes our techniques robust and versatile. Experiments show that our techniques bring a notable improvement on efficiency and effectiveness.