2025/05/23 by Ananth Muppidi, Tarak Das, Muppidi, Ananth +7
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Benchmark (surveying) #Computation and Language (cs.CL) #Content (measure theory) #FOS: Computer and information sciences #Generative grammar #Key (lock) #Natural Language Processing Techniques #Presentation (obstetrics) #Set (abstract data type)
paper · pdf · doi:10.48550/arxiv.2505.18240
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The generation of presentation slides automatically is an important problem in the era of generative AI. This paper focuses on evaluating multimodal content in presentation slides that can effectively summarize a document and convey concepts to a broad audience. We introduce a benchmark dataset, RefSlides, consisting of human-made high-quality presentations that span various topics. Next, we propose a set of metrics to characterize different intrinsic properties of the content of a presentation and present REFLEX, an evaluation approach that generates scores and actionable feedback for these metrics. We achieve this by generating negative presentation samples with different degrees of metric-specific perturbations and use them to fine-tune LLMs. This reference-free evaluation technique does not require ground truth presentations during inference. Our extensive automated and human experiments demonstrate that our evaluation approach outperforms classical heuristic-based and state-of-the-art large language model-based evaluations in generating scores and explanations.