2020/12/07 by David Sun, Hadas Kotek, Sun, David Q. +9
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Mobile Crowdsensing and Crowdsourcing #Natural Language Processing Techniques #Software Engineering Research
paper · pdf · doi:10.48550/arxiv.2012.04169
openalex publication_date 2020/12/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper develops and implements a scalable methodology for (a) estimating\nthe noisiness of labels produced by a typical crowdsourcing semantic annotation\ntask, and (b) reducing the resulting error of the labeling process by as much\nas 20-30% in comparison to other common labeling strategies. Importantly, this\nnew approach to the labeling process, which we name Dynamic Automatic Conflict\nResolution (DACR), does not require a ground truth dataset and is instead based\non inter-project annotation inconsistencies. This makes DACR not only more\naccurate but also available to a broad range of labeling tasks. In what follows\nwe present results from a text classification task performed at scale for a\ncommercial personal assistant, and evaluate the inherent ambiguity uncovered by\nthis annotation strategy as compared to other common labeling strategies.\n