vix.ing · top · new · best · stats · spec

Time-Sensitive Bayesian Information Aggregation for Crowdsourcing\n Systems

2015/10/21 by Matteo Venanzi, Venanzi, Matteo, John Guiver +4 · 1 citation
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Human Mobility and Location-Based Analysis #Machine Learning (cs.LG) #Mobile Crowdsensing and Crowdsourcing #Transportation Planning and Optimization

paper · pdf · doi:10.48550/arxiv.1510.06335

openalex publication_date 2015/10/21 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28

Abstract

Crowdsourcing systems commonly face the problem of aggregating multiple\njudgments provided by potentially unreliable workers. In addition, several\naspects of the design of efficient crowdsourcing processes, such as defining\nworker's bonuses, fair prices and time limits of the tasks, involve knowledge\nof the likely duration of the task at hand. Bringing this together, in this\nwork we introduce a new time--sensitive Bayesian aggregation method that\nsimultaneously estimates a task's duration and obtains reliable aggregations of\ncrowdsourced judgments. Our method, called BCCTime, builds on the key insight\nthat the time taken by a worker to perform a task is an important indicator of\nthe likely quality of the produced judgment. To capture this, BCCTime uses\nlatent variables to represent the uncertainty about the workers' completion\ntime, the tasks' duration and the workers' accuracy. To relate the quality of a\njudgment to the time a worker spends on a task, our model assumes that each\ntask is completed within a latent time window within which all workers with a\npropensity to genuinely attempt the labelling task (i.e., no spammers) are\nexpected to submit their judgments. In contrast, workers with a lower\npropensity to valid labeling, such as spammers, bots or lazy labelers, are\nassumed to perform tasks considerably faster or slower than the time required\nby normal workers. Specifically, we use efficient message-passing Bayesian\ninference to learn approximate posterior probabilities of (i) the confusion\nmatrix of each worker, (ii) the propensity to valid labeling of each worker,\n(iii) the unbiased duration of each task and (iv) the true label of each task.\nUsing two real-world public datasets for entity linking tasks, we show that\nBCCTime produces up to 11% more accurate classifications and up to 100% more\ninformative estimates of a task's duration compared to state-of-the-art\nmethods.\n

Citations

Cited by

Related