Free Lunch for User Experience: Crowdsourcing Agents for Scalable User Studies
2025/05/29 by Siyang Liu, Sahand Sabour, Liu, Siyang +5 · 1 citation
Computer Science · Social Sciences · #Mobile Crowdsensing and Crowdsourcing #Spreadsheets and End-User Computing #Ethics and Social Impacts of AI
paper · pdf · doi:10.48550/arxiv.2505.22981
Abstract
User studies are central to user experience research, yet recruiting participant is expensive, slow, and limited in diversity. Recent work has explored using Large Language Models as simulated users, but doubts about fidelity have hindered practical adoption. We deepen this line of research by asking whether scale itself can enable useful simulation, even if not perfectly accurate. We introduce Crowdsourcing Simulated User Agents, a method that recruits generative agents from billion-scale profile assets to act as study participants. Unlike handcrafted simulations, agents are treated as recruitable, screenable, and engageable across UX research stages. To ground this method, we demonstrate a game prototyping study with hundreds of simulated players, comparing their insights against a 10-participant local user study and a 20-participant crowdsourcing study with humans. We find a clear scaling effect: as the number of simulated user agents increases, coverage of human findings rises smoothly and plateaus around 90%. 12.8 simulated agents are as useful as one locally recruited human, and 3.2 agents are as useful as one crowdsourced human. Results show that while individual agents are imperfect, aggregated simulations produce representative and actionable insights comparable to real users. Professional designers further rated these insights as balancing fidelity, cost, time efficiency, and usefulness. Finally, we release an agent crowdsourcing toolkit with a modular open-source pipeline and a curated pool of profiles synced from ongoing simulation research, to lower the barrier for researchers to adopt simulated participants. Together, this work contributes a validated method and reusable toolkit that expand the options for conducting scalable and practical UX studies.
Citations
- Mind the (Belief) Gap: Group Identity in the World of LLMs
- Eeyore: Realistic Depression Simulation via Supervised and Preference Optimization
- Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas
- TeachTune: Reviewing Pedagogical Agents Against Diverse Student Profiles with Simulated Students
- 'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants
- Simulated patient systems powered by large language model-based AI agents offer potential for transforming medical education
- Reclaiming AI as a Theoretical Tool for Cognitive Science
- LLMs generate structurally realistic social networks but overestimate political homophily
- Proxona: Supporting Creators' Sensemaking and Ideation with LLM-Powered Audience Personas
- Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas
- Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- PATIENT-Ψ: Using Large Language Models to Simulate Patients for Training Mental Health Professionals
- From Role-Play to Drama-Interaction: An LLM Solution
- From Persona to Personalization: A Survey on Role-Playing Language Agents
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- AI for social science and social science of AI: A Survey
- You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
- LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature Engineering
- Generative Agents: Interactive Simulacra of Human Behavior
- Data quality in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA
- The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development
- Out of One, Many: Using Language Models to Simulate Human Samples
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Towards Process-Oriented, Modular, and Versatile Question Generation that Meets Educational Needs
- Personalizing Dialogue Agents: I have a dog, do you have pets too?
- From Factors to Actors: Computational Sociology and Agent-Based Modeling
- THE DISTRIBUTION OF THE FLORA IN THE ALPINE ZONE. 1
Cited by
Related