2024/03/18 by Yifan Ding, Ding, Yifan, Michael Yankoski +3 · 2 citations
Computer Science · Decision Sciences · #Advanced Database Systems and Queries #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2403.15453
openalex publication_date 2024/03/18 · openalex created_date 2024/03/27 · openalex updated_date 2026/07/28
Information Extraction refers to a collection of tasks within Natural Language Processing (NLP) that identifies sub-sequences within text and their labels. These tasks have been used for many years to link extract relevant information and to link free text to structured data. However, the heterogeneity among information extraction tasks impedes progress in this area. We therefore offer a unifying perspective centered on what we define to be spans in text. We then re-orient these seemingly incongruous tasks into this unified perspective and then re-present the wide assortment of information extraction tasks as variants of the same basic Span-Oriented Information Extraction task.