VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
2025/06/03 by Xu, Zelai, Xu, Zhexuan, Yi, Xiangmin +7 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2506.02387
Abstract
Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain limited to single-agent or text-only environments. In contrast, real-world scenarios often involve multiple agents interacting within rich visual and textual contexts, posing challenges with both multimodal observations and strategic interactions. To bridge this gap, we introduce Visual Strategic Bench (VS-Bench), a multimodal benchmark that evaluates VLMs for strategic abilities in multi-agent environments. VS-Bench comprises ten vision-grounded environments that cover cooperative, competitive, and mixed-motive interactions. The performance of VLM agents is evaluated across three dimensions: perception measured by element recognition accuracy; strategic reasoning measured by next-action prediction accuracy; and decision-making measured by normalized episode return. Extensive experiments on fifteen leading VLMs show that, although current models exhibit strong perception abilities, there remains a significant gap to optimal performance in reasoning and decision-making, with the best-performing model attaining 46.6% prediction accuracy and 31.4% normalized return. We further analyze the key factors influencing performance, conduct human experiments, and examine failure modes to provide a deeper understanding of VLMs' strategic abilities. By standardizing the evaluation and highlighting the limitations of existing models, we envision VS-Bench as a foundation for future research on strategic multimodal agents. Code and data are available at https://vs-bench.github.io.
Citations
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Are Large Vision Language Models Good Game Players?
- Qwen2.5-VL Technical Report
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
- How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
- On the Effects of Data Scale on UI Control Agents
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
- LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
- GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- GPT-4V(ision) is a Generalist Web Agent, if Grounded
- Creative Agents: Empowering Agents with Imagination for Creative Tasks
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
- MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
- LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- ProAgent: Building Proactive Cooperative Agents with Large Language Models
- Visual Instruction Tuning
- Language Instructed Reinforcement Learning for Human-AI Coordination
- LLaMA: Open and Efficient Foundation Language Models
- Model-Free Opponent Shaping
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Collaborating with Humans without Human Data
- Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
- The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games
- PettingZoo: Gym for Multi-Agent Reinforcement Learning
- "Other-Play" for Zero-Shot Coordination
- Dota 2 with Large Scale Deep Reinforcement Learning
- On the Utility of Learning about Humans for Human-AI Coordination
- A Generalized Training Approach for Multiagent Learning
- OpenSpiel: A Framework for Reinforcement Learning in Games
- The StarCraft Multi-Agent Challenge
- The Hanabi Challenge: A New Frontier for AI Research
- Learning with Opponent-Learning Awareness
- Prosocial learning agents solve generalized Stag Hunts better than selfish ones
- Maintaining cooperation in complex social dilemmas using deep reinforcement learning
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- VQA: Visual Question Answering
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Playing Atari with Deep Reinforcement Learning
- Bayes' Bluff: Opponent Modelling in Poker
- Gemini Robotics: Bringing AI into the Physical World
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Cited by
Related