2024/08/17 by Sher Badshah, Badshah, Sher, Hassan Sajjad +1 · 15 citations
Social Sciences · #68T07 #68T20 #68T50 #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Comparative and International Law Studies #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.0 #I.2.2 #I.2.7 #Legal Education and Practice Innovations
paper · pdf · doi:10.48550/arxiv.2408.09235
openalex publication_date 2024/08/17 · openalex created_date 2024/10/01 · openalex updated_date 2026/07/28
The emergence of Large Language Models (LLMs) as chat assistants capable of generating human-like conversations has amplified the need for robust evaluation methods, particularly for open-ended tasks. Conventional metrics such as EM and F1, while useful, are inadequate for capturing the full semantics and contextual depth of such generative outputs. We propose a reference-guided verdict method that automates the evaluation process by leveraging multiple LLMs as judges. Through experiments on free-form question-answering tasks, we demonstrate that combining multiple models improves the reliability and accuracy of evaluations, especially in tasks where a single model may struggle. The results indicate a strong correlation with human evaluations, establishing the proposed method as a reliable alternative to traditional metrics.