2025/04/11 by Leo Kampen, Carlos Rabat Villarreal, Kampen, Leo +7 · 1 citation
Social Sciences · #Artificial Intelligence (cs.AI) #Asian Studies and History #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2504.08211
openalex publication_date 2025/04/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we conducted a Multi-Perspective Comparative Narrative Analysis (CNA) on three prominent LLMs: GPT-3.5, PaLM2, and Llama2. We applied identical prompts and evaluated their outputs on specific tasks, ensuring an equitable and unbiased comparison between various LLMs. Our study revealed that the three LLMs generated divergent responses to the same prompt, indicating notable discrepancies in their ability to comprehend and analyze the given task. Human evaluation was used as the gold standard, evaluating four perspectives to analyze differences in LLM performance.