vix.ing · top · new · best · stats · spec

Easy, Reproducible and Quality-Controlled Data Collection with Crowdaq

2020/10/06 by Qiang Ning, Ning, Qiang, Hao Wu +13
Computer Science · Decision Sciences · #Data Stream Mining Techniques #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Mobile Crowdsensing and Crowdsourcing #Scientific Computing and Data Management

paper · pdf · doi:10.48550/arxiv.2010.06694

openalex publication_date 2020/10/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

High-quality and large-scale data are key to success for AI systems. However, large-scale data annotation efforts are often confronted with a set of common challenges: (1) designing a user-friendly annotation interface; (2) training enough annotators efficiently; and (3) reproducibility. To address these problems, we introduce Crowdaq, an open-source platform that standardizes the data collection pipeline with customizable user-interface components, automated annotator qualification, and saved pipelines in a re-usable format. We show that Crowdaq simplifies data annotation significantly on a diverse set of data collection use cases and we hope it will be a convenient tool for the community.

Citations

Related