2023/09/18 by Alejandro Cuevas, Cuevas, Alejandro, Jennifer V. Scurrell +8 · 2 voices · 4 citations
Computer Science · Psychology · Social Sciences · #AI in Service Interactions #Applied psychology #Chatbot #Computer science #Data science #Expert finding and Q&A systems #Interview #Machine learning #Psychology #Qualitative property #Quality (philosophy) #Scale (ratio) #Social Media and Politics #World Wide Web
paper · pdf · doi:10.48550/arxiv.2309.10187
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/18 · openalex created_date 2023/09/21 · openalex updated_date 2026/07/28
Chatbots have shown promise as tools to scale qualitative data collection. Recent advances in Large Language Models (LLMs) could accelerate this process by allowing researchers to easily deploy sophisticated interviewing chatbots. We test this assumption by conducting a large-scale user study (n=399) evaluating 3 different chatbots, two of which are LLM-based and a baseline which employs hard-coded questions. We evaluate the results with respect to participant engagement and experience, established metrics of chatbot quality grounded in theories of effective communication, and a novel scale evaluating "richness" or the extent to which responses capture the complexity and specificity of the social context under study. We find that, while the chatbots were able to elicit high-quality responses based on established evaluation metrics, the responses rarely capture participants' specific motives or personalized examples, and thus perform poorly with respect to richness. We further find low inter-rater reliability between LLMs and humans in the assessment of both quality and richness metrics. Our study offers a cautionary tale for scaling and evaluating qualitative research with LLMs.