vix.ing · top · new · best · stats

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning

2025/08/29 by João Guilherme Alves Santos, Santos, João Guilherme Alves, Giovana Kerche Bonás +3 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Benchmark (surveying) #Closed captioning #Computation and Language (cs.CL) #Context (archaeology) #FOS: Computer and information sciences #Leverage (statistics) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Relevance (law) #Text Readability and Simplification #USable

paper · pdf · doi:10.48550/arxiv.2508.21294

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/08/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX dataset, now including 2024-2025 exams and automatically generated image captions using state-of-the-art models, enhancing its relevance for data contamination studies in LLM pretraining. Captioning strategies increase accessibility to text-only models by more than 40%, producing 1,422 usable questions, more than doubling the number in the original BLUEX. We evaluated commercial and open-source LLMs and their ability to leverage visual context through captions.

Citations

Cited by

Related