2026/04/07 by Wang Yang, Chaoda Song, Xinpeng Li +7 · 1 voice
Computer Science · #AI-based Problem Solving and Planning #Benchmark (surveying) #Consistency (knowledge bases) #Decoy #Domain (mathematical analysis) #Multimodal Machine Learning Applications #Overhead (engineering) #Reinforcement Learning in Robotics #Scalability #Schedule #Task (project management) #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2604.06111
openalex publication_date 2026/04/07 · arxiv published 2026/04/07 · openalex created_date 2026/04/09 · arxiv updated 2026/04/10 · openalex updated_date 2026/07/28
Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41% of total evaluation time) and imbalanced task horizon and difficulty distributions that make aggregate scores unreliable. To address these issues, we propose AgentCE-Bench built around a unified grid-based planning task, where agents must fill hidden slots in a partially completed schedule subject to both local slot constraints and global constraints. Our benchmark offers fine-grained control through two orthogonal axes: Scalable Horizons, controlled by the number of hidden slots H, and Controllable Difficulty, governed by a decoy budget B that determines the number of globally misleading decoy candidates. Crucially, all tool calls are resolved via static JSON files under a Lightweight Environment design, eliminating setup overhead and enabling fast, reproducible evaluation suitable for training-time validation. We first validate that H and B provide reliable control over task horizon and difficulty, and that AgentCE-Bench exhibits strong domain consistency and model discriminability. We then conduct comprehensive experiments across 13 models of diverse sizes and families over 6 domains, revealing significant cross-model performance variation and confirming that AgentCE-Bench provides interpretable and controllable evaluation of agent reasoning.