2021/11/17 by S. Vamshi Krishna, Singla, Yaman Kumar, Krishna, Sriram +4 · 1 citation
Computer Science · #Applications (stat.AP) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Testing and Debugging Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2111.08906
openalex publication_date 2021/11/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Automated Scoring (AS), the natural language processing task of scoring\nessays and speeches in an educational testing setting, is growing in popularity\nand being deployed across contexts from government examinations to companies\nproviding language proficiency services. However, existing systems either forgo\nhuman raters entirely, thus harming the reliability of the test, or score every\nresponse by both human and machine thereby increasing costs. We target the\nspectrum of possible solutions in between, making use of both humans and\nmachines to provide a higher quality test while keeping costs reasonable to\ndemocratize access to AS. In this work, we propose a combination of the\nexisting paradigms, sampling responses to be scored by humans intelligently. We\npropose reward sampling and observe significant gains in accuracy (19.80%\nincrease on average) and quadratic weighted kappa (QWK) (25.60% on average)\nwith a relatively small human budget (30% samples) using our proposed sampling.\nThe accuracy increase observed using standard random and importance sampling\nbaselines are 8.6% and 12.2% respectively. Furthermore, we demonstrate the\nsystem's model agnostic nature by measuring its performance on a variety of\nmodels currently deployed in an AS setting as well as pseudo models. Finally,\nwe propose an algorithm to estimate the accuracy/QWK with statistical\nguarantees (Our code is available at https://git.io/J1IOy).\n