2021/06/02 by Paul Resnick, Yuqing Kong, Resnick, Paul +5 · 2 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Misinformation and Its Impacts #Multiagent Systems (cs.MA) #Spam and Phishing Detection
paper · pdf · doi:10.48550/arxiv.2106.01254
openalex publication_date 2021/06/02 · openalex created_date 2021/06/22 · openalex updated_date 2026/07/28
In many classification tasks, there is no definitive ground truth, only human judgments that may disagree. We address two challenges that arise in such settings: (1) how to use human raters to score classifiers, and (2) how to use them for comparison benchmarks. For the first, the common practice is to score classifiers against the majority vote of an evaluation panel of several human raters. We argue that this is not justified when either of two properties fails: objectivity or equanimity. Instead, under a utility model appropriate for such settings, scoring against one rater at a time and averaging the scores across raters is a more principled approach. For the second, we introduce the concept of rater equivalence: the smallest number of human raters whose combined judgment matches the classifier's performance. We provide a provably optimal algorithm for combining benchmark panel labels, and demonstrate the framework through case studies.