vix.ing · top · new · best · stats

Graph integration of structured, semistructured and unstructured data for data journalism

2020/12/16 by Angelos-Christos G. Anadiotis, Angelos-Christos Anadiotis, Oana Balalau +16
Computer Science · Decision Sciences · #Advanced Database Systems and Queries #Artificial Intelligence (cs.AI) #Databases (cs.DB) #FOS: Computer and information sciences #Scientific Computing and Data Management #Web Data Mining and Analysis #cs.AI #cs.DB

paper · pdf · doi:10.48550/arxiv.2012.08830

40 pages, 9 figures. arXiv admin note: substantial text overlap with arXiv:2007.12488, arXiv:2009.04283

arxiv created 2020/12/16 · openalex publication_date 2020/12/16 · arxiv updated 2020/12/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Digital data is a gold mine for modern journalism. However, datasets which interest journalists are extremely heterogeneous, ranging from highly structured (relational databases), semi-structured (JSON, XML, HTML), graphs (e.g., RDF), and text. Journalists (and other classes of users lacking advanced IT expertise, such as most non-governmental-organizations, or small public administrations) need to be able to make sense of such heterogeneous corpora, even if they lack the ability to define and deploy custom extract-transform-load workflows, especially for dynamically varying sets of data sources. We describe a complete approach for integrating dynamic sets of heterogeneous datasets along the lines described above: the challenges we faced to make such graphs useful, allow their integration to scale, and the solutions we proposed for these problems. Our approach is implemented within the ConnectionLens system; we validate it through a set of experiments.

Citations

Related