vix.ing · top · new · best · stats

VLMs for Videogame Data Annotation

2026/08/06 by Katrin Schmid, Iuri Frosio · 1 citation
Computer Science · #cs.AI #cs.CV #cs.LG

paper · pdf

arxiv created 2026/08/06 · arxiv updated 2026/08/07

Abstract

Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.

Citations

Cited by