2021/01/16 by Nicholas Clark, Matthew Dabkowski, Patrick Driscoll +4
Business, Management and Accounting · Computer Science · Engineering · Mathematics · #Artificial intelligence #Computer science #Confidence interval #Credibility #Data mining #Econometrics #Geography #Green IT and Sustainability #Leverage (statistics) #Mathematics #Quality Function Deployment in Product Design #Sample (material) #Sample size determination #Scale (ratio) #Software Reliability and Analysis Research #Statistics #Usability #stat.ME
paper · pdf · doi:10.1080/10447318.2020.1870831
arxiv created 2021/01/16 · openalex publication_date 2021/01/25 · arxiv updated 2021/01/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The System Usability Scale (SUS) is a short, survey-based approach used to determine the usability of a system from an end-user perspective once a prototype is available for assessment. Individual scores are gathered using a ten-question survey with the survey results reported in terms of central tendency (sample mean) as an estimate of the system’s usability (the SUS study score), and confidence intervals (CIs) on the sample mean are used to communicate uncertainty levels associated with this point estimate. When the number of individuals surveyed is large, the SUS study scores and accompanying confidence intervals relying upon the central limit theorem for support are appropriate. However, when only a small number of users are surveyed, reliance on the central limit theorem falls short, resulting in CIs that suffer from parameter bound violations and interval widths that confound mappings to adjective and other constructed scales. These shortcomings are especially pronounced when the underlying SUS score data is skewed, as it is in many instances. This paper introduces an empirically based remedy for such small-sample circumstances, proposing a set of decision rules that leverage either an extended bias-corrected accelerated (BCa) bootstrap confidence interval (Cl) or an empirical Bayesian credibility interval about the sample mean to restore and bolster subsequent Cl accuracy. Data from historical SUS assessments are used to highlight shortfalls in current practices and to demonstrate the improvements these alternate approaches offer while remaining statistically defensible. A freely available, online application is introduced and discussed that automates SUS analysis under these decision rules, thereby assisting usability practitioners in adopting the advocated approaches.