vix.ing · top · new · best · stats · spec

Automatic vs Manual Provenance Abstractions: Mind the Gap

2016/05/21 by Pinar Alper, Alper, Pinar, Khalid Belhajjame +4
Computer Science · Decision Sciences · #Abstraction #Artificial intelligence #Computer science #Context (archaeology) #Data science #Database #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Filter (signal processing) #Human–computer interaction #Perspective (graphical) #Process (computing) #Programming language #Research Data Management Practices #Scientific Computing and Data Management #Software Engineering (cs.SE) #Software engineering #Workflow #Workflow engine #Workflow technology #Zoom #cs.SE

paper · pdf · doi:10.48550/arxiv.1605.06669

Preprint accepted to the 2016 workshop on the Theory and Applications of Provenance, TAPP 2016

arxiv created 2016/05/21 · openalex publication_date 2016/05/21 · arxiv updated 2016/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In recent years the need to simplify or to hide sensitive information in provenance has given way to research on provenance abstraction. In the context of scientific workflows, existing research provides techniques to semi automatically create abstractions of a given workflow description, which is in turn used as filters over the workflow's provenance traces. An alternative approach that is commonly adopted by scientists is to build workflows with abstractions embedded into the workflow's design, such as using sub-workflows. This paper reports on the comparison of manual versus semi-automated approaches in a context where result abstractions are used to filter report-worthy results of computational scientific analyses. Specifically; we take a real-world workflow containing user-created design abstractions and compare these with abstractions created by ZOOM UserViews and Workflow Summaries systems. Our comparison shows that semi-automatic and manual approaches largely overlap from a process perspective, meanwhile, there is a dramatic mismatch in terms of data artefacts retained in an abstracted account of derivation. We discuss reasons and suggest future research directions.

Related