2016/03/25 by Karthik Narasimhan, Adam Yala, Narasimhan, Karthik +3 · 5 citations
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Topic Modeling #Web Data Mining and Analysis
paper · pdf · doi:10.48550/arxiv.1603.07954
openalex publication_date 2016/03/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Most successful information extraction systems operate with access to a large\ncollection of documents. In this work, we explore the task of acquiring and\nincorporating external evidence to improve extraction accuracy in domains where\nthe amount of training data is scarce. This process entails issuing search\nqueries, extraction from new sources and reconciliation of extracted values,\nwhich are repeated until sufficient evidence is collected. We approach the\nproblem using a reinforcement learning framework where our model learns to\nselect optimal actions based on contextual information. We employ a deep\nQ-network, trained to optimize a reward function that reflects extraction\naccuracy while penalizing extra effort. Our experiments on two databases -- of\nshooting incidents, and food adulteration cases -- demonstrate that our system\nsignificantly outperforms traditional extractors and a competitive\nmeta-classifier baseline.\n