vix.ing · top · new · best · stats · spec

Towards Standard Criteria for human evaluation of Chatbots: A Survey

2021/05/24 by Hongru Liang, Huaqing Li, Liang, Hongru +1
Computer Science · Decision Sciences · #AI in Service Interactions #Computation and Language (cs.CL) #FOS: Computer and information sciences #Personal Information Management and User Behavior #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2105.11197

openalex publication_date 2021/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Human evaluation is becoming a necessity to test the performance of Chatbots. However, off-the-shelf settings suffer the severe reliability and replication issues partly because of the extremely high diversity of criteria. It is high time to come up with standard criteria and exact definitions. To this end, we conduct a through investigation of 105 papers involving human evaluation for Chatbots. Deriving from this, we propose five standard criteria along with precise definitions.

Citations

Related