2025/03/27 by Ashish Sardana, Sardana, Ashish · 1 voice · 2 citations
Computer Science · Medicine · Neuroscience · #Epilepsy research and treatment #FOS: Computer and information sciences #Functional Brain Connectivity Studies #Machine Learning (cs.LG) #Treatment of Major Depression #cs.LG
paper · pdf · doi:10.48550/arxiv.2503.21157
openalex publication_date 2025/03/27 · arxiv published 2025/03/27 · arxiv updated 2025/04/07 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28
This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods included in our study include: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination Evaluation Model (HHEM), and the Trustworthy Language Model (TLM). These approaches are all reference-free, requiring no ground-truth answers/labels to catch incorrect LLM responses. Our study reveals that, across diverse RAG applications, some of these approaches consistently detect incorrect RAG responses with high precision/recall.