vix.ing · top · new · best · stats · spec

Towards Interpretable and Learnable Risk Analysis for Entity Resolution

2019/12/06 by Zhaoqiang Chen, Chen, Zhaoqiang, Qun Chen +9 · 1 citation
Computer Science · Decision Sciences · #Data Mining Algorithms and Applications #Data Quality and Management #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1912.02947

openalex publication_date 2019/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Machine-learning-based entity resolution has been widely studied. However, some entity pairs may be mislabeled by machine learning models and existing studies do not study the risk analysis problem -- predicting and interpreting which entity pairs are mislabeled. In this paper, we propose an interpretable and learnable framework for risk analysis, which aims to rank the labeled pairs based on their risks of being mislabeled. We first describe how to automatically generate interpretable risk features, and then present a learnable risk model and its training technique. Finally, we empirically evaluate the performance of the proposed approach on real data. Our extensive experiments have shown that the learning risk model can identify the mislabeled pairs with considerably higher accuracy than the existing alternatives.

Citations

Cited by

Related