2024/03/28 by Mubashara Akhtar, Omar Benjelloun, Costanza Conforti +30 · 1 voice · 3 citations
Computer Science · Decision Sciences · #Advanced Data Storage Technologies #Machine Learning and Data Classification #Scientific Computing and Data Management #cs.AI #cs.DB #cs.IR #cs.LG
paper · pdf · doi:10.1145/3650203.3663326
arxiv published 2024/03/28 · openalex created_date 2024/03/30 · openalex publication_date 2024/05/29 · arxiv updated 2024/12/09 · openalex updated_date 2026/08/04
Data is a critical resource for Machine Learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that simplifies how data is used by ML tools and frameworks. Croissant makes datasets more discoverable, portable and interoperable, thereby addressing significant challenges in ML data management and responsible AI. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, ready to be loaded into the most popular ML frameworks.