2013/11/14 by Seyed-Mehdi-Reza Beheshti, Srikumar Venugopal, Beheshti, Seyed-Mehdi-Reza +7 · 24 citations
Computer Science · Decision Sciences · #Artificial intelligence #Big data #Byte #Computation and Language (cs.CL) #Computer science #Coreference #Data Quality and Management #Data mining #Data science #Distributed #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Information retrieval #Natural Language Processing Techniques #Parallel #Process (computing) #Resolution (logic) #State (computer science) #Task (project management) #Topic Modeling #and Cluster Computing (cs.DC) #cs.CL #cs.DC #cs.IR
paper · pdf · doi:10.48550/arxiv.1311.3987
published in arXiv (Cornell University) (Cornell University)
arxiv created 2013/11/14 · openalex publication_date 2013/11/14 · arxiv updated 2013/11/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Information Extraction (IE) is the task of automatically extracting structured information from unstructured/semi-structured machine-readable documents. Among various IE tasks, extracting actionable intelligence from ever-increasing amount of data depends critically upon Cross-Document Coreference Resolution (CDCR) - the task of identifying entity mentions across multiple documents that refer to the same underlying entity. Recently, document datasets of the order of peta-/tera-bytes has raised many challenges for performing effective CDCR such as scaling to large numbers of mentions and limited representational power. The problem of analysing such datasets is called "big data". The aim of this paper is to provide readers with an understanding of the central concepts, subtasks, and the current state-of-the-art in CDCR process. We provide assessment of existing tools/techniques for CDCR subtasks and highlight big data challenges in each of them to help readers identify important and outstanding issues for further investigation. Finally, we provide concluding remarks and discuss possible directions for future work.