2017/06/27 by Kevin Jasberg, Jasberg, Kevin, Sergej Sizov +1 · 1 citation
Decision Sciences · #Multi-Criteria Decision Making
paper · pdf · doi:10.48550/arxiv.1706.08866
In this paper, we examine the statistical soundness of comparative\nassessments within the field of recommender systems in terms of reliability and\nhuman uncertainty. From a controlled experiment, we get the insight that users\nprovide different ratings on same items when repeatedly asked. This volatility\nof user ratings justifies the assumption of using probability densities instead\nof single rating scores. As a consequence, the well-known accuracy metrics\n(e.g. MAE, MSE, RMSE) yield a density themselves that emerges from convolution\nof all rating densities. When two different systems produce different RMSE\ndistributions with significant intersection, then there exists a probability of\nerror for each possible ranking. As an application, we examine possible ranking\nerrors of the Netflix Prize. We are able to show that all top rankings are more\nor less subject to high probabilities of error and that some rankings may be\ndeemed to be caused by mere chance rather than system quality.\n