vix.ing · top · new · best · stats · spec

Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?

2025/03/27 by Ashish Sardana, Sardana, Ashish · 1 voice · 2 citations
Computer Science · Medicine · Neuroscience · #Epilepsy research and treatment #FOS: Computer and information sciences #Functional Brain Connectivity Studies #Machine Learning (cs.LG) #Treatment of Major Depression #cs.LG

paper · pdf · doi:10.48550/arxiv.2503.21157

openalex publication_date 2025/03/27 · arxiv published 2025/03/27 · arxiv updated 2025/04/07 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28

Abstract

This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods included in our study include: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination Evaluation Model (HHEM), and the Trustworthy Language Model (TLM). These approaches are all reference-free, requiring no ground-truth answers/labels to catch incorrect LLM responses. Our study reveals that, across diverse RAG applications, some of these approaches consistently detect incorrect RAG responses with high precision/recall.

Cited by

Discussions

Related