2025/08/09 by Hyo Jin Do, Do, Hyo Jin, Rachel Ostrand +10 · 1 voice
Computer Science · Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Baseline (sea) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Hallucinating #Human-Computer Interaction (cs.HC) #Readability #Style (visual arts) #Transparency (behavior) #Usability #cs.AI #cs.HC
paper · pdf · doi:10.48550/arxiv.2508.06846
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/08/09 · arxiv published 2025/08/09 · arxiv updated 2025/08/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advancements have been made to detect hallucinated content by assessing the factuality of the model's responses, there is still limited research on how to effectively communicate this information to users. To address this gap, we conducted two scenario-based experiments with a total of 208 participants to systematically compare the effects of various design strategies for communicating factuality scores by assessing participants' ratings of trust, ease in validating response accuracy, and preference. Our findings reveal that participants preferred and trusted a design in which all phrases within a response were color-coded based on factuality scores. Participants also found it easier to validate accuracy of the response in this style compared to a baseline with no style applied. Our study offers practical design guidelines for LLM application developers and designers, aimed at calibrating user trust, aligning with user preferences, and enhancing users' ability to scrutinize LLM outputs.