2016/10/13 by Muhammad Rezaul Karim, Karim, Muhammad Rezaul, David W. Messinger +6
Computer Science · #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Mobile Crowdsensing and Crowdsourcing #Open Source Software Innovations #Software Engineering (cs.SE) #Software Engineering Research
paper · pdf · doi:10.48550/arxiv.1610.04142
openalex publication_date 2016/10/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Crowdsourced software development (CSD) offers a series of specified tasks to\na large crowd of trustworthy software workers. Topcoder is a leading platform\nto manage the whole process of CSD. While increasingly accepted as a realistic\noption for software development, preliminary analysis on Topcoder's software\ncrowd worker behaviors reveals an alarming task-quitting rate of 82.9%. In\naddition, a substantial number of tasks do not receive any successful\nsubmission.\n In this paper, we report about a methodology to improve the efficiency of\nCSD. We apply massive data analytics and machine leaning to (i) perform\ncomparative analysis on alternative technique analysis to predict likelihood of\nwinners and quitters for each task, (ii) significantly reduce the amount of\nnon-succeeding development effort in registered but inappropriate tasks, (iii)\nidentify and rank the most qualified registered workers for each task, and (iv)\nprovide reliable prediction of tasks risky to get any successful submission.\n Our results and analysis show that Random Forest (RF) based predictive\ntechnique performs best among the alternative techniques studied. Applying RF,\nthe tasks recommended to workers can reduce the amount of non-succeeding\ndevelopment effort to a great extent. On average, over a period of 30 days, the\nsavings are 3.5 and 4.6 person-days per registered tasks for experienced resp.\nunexperienced workers. For the task-related recommendations of workers, we can\naccurately recommend at least 1 actual winner in the top ranked workers,\nparticularly 94.07% of the time among the top-2 recommended workers for each\ntask. Finally, we can predict, with more than 80% F-measure, the tasks likely\nnot getting any submission, thus triggering timely corrective actions from CSD\nplatforms or task requesters.\n