vix.ing · top · new · best · stats · spec

Retrieve, Merge, Predict: Augmenting Tables with Data Lakes

2024/01/31 by Riccardo Cappuzzo, Cappuzzo, Riccardo, Coelho, Aimee +5 · 2 citations
Computer Science · Decision Sciences · #Advanced Database Systems and Queries #Data Mining Algorithms and Applications #Data Quality and Management #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.2402.06282

openalex publication_date 2024/01/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)" We present an in-depth analysis of data discovery for analytics in data lakes, focusing on table augmentation for given machine learning tasks. We analyze alternative methods used in the three key steps: retrieving joinable tables, merging information, and predicting with the resultant table. As data lakes, the paper uses YADL (Yet Another Data Lake) -- a novel dataset developed as a tool for benchmarking this data discovery task -- and Open Data US, a well-referenced real data lake. Through systematic exploration on both lakes, our study outlines the importance of accurately retrieving join candidates, and the efficiency of simple aggregation methods. We report new insights on the benefits of existing solutions and on the their limitations, aiming at guiding future research in this space.

Cited by

Related