vix.ing · top · new · best · stats · spec

A Scalable Framework for Quality Assessment of RDF Datasets

2019/10/30 by Gëzim Sejdiu, Anisa Rula, Sejdiu, Gezim +5
Computer Science · #Advanced Database Systems and Queries #Databases (cs.DB) #FOS: Computer and information sciences #Performance (cs.PF) #Semantic Web and Ontologies #Service-Oriented Architecture and Web Services

paper · pdf · doi:10.48550/arxiv.2001.11100

openalex publication_date 2019/10/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Over the last years, Linked Data has grown continuously. Today, we<br> than 10,000 datasets being available online following Linked Data<br> These standards allow data to be machine readable and inter-operable.<br> ss, many applications, such as data integration, search, and interlink-<br> take full advantage of Linked Data if it is of low quality. There exist a<br> ches for the quality assessment of Linked Data, but their performance<br> ith the increase in data size and quickly grows beyond the capabilities<br> machine. In this paper, we present DistQualityAssessment – an open<br> lementation of quality assessment of large RDF datasets that can scale<br> ster of machines. This is the first distributed, in-memory approach for<br> different quality metrics for large RDF datasets using Apache Spark.<br> ovide a quality assessment pattern that can be used to generate new<br> etrics that can be applied to big data. The work presented here is in-<br> th the SANSA framework and has been applied to at least three use<br> nd the SANSA community. The results show that our approach is more<br> icient, and scalable as compared to previously proposed approaches.<br>

Related