2026/04/04 by Shaoyang Cui, Lingbei Meng, Yaodi Luo +1
Computer Science · #cs.CV
7 pages, 5 figures
arxiv created 2026/07/29 · arxiv updated 2026/07/30
Video-grounded numerical reasoning requires Vision-Language Models (VLMs) to identify, track, and combine quantitative evidence across frames, actions, and scene changes. Existing benchmarks provide fragmented coverage: general VideoQA includes counting among broader tasks, while dedicated benchmarks focus on repetition counting, ultra-long-video enumeration, or instructional mathematics. We introduce VidNum, a manually curated and independently verified benchmark containing 1,167 multiple-choice questions. Its three task groups distinguish Direct and Distinct Enumeration, Conditioned and Structured Enumeration, and Compositional Quantitative Reasoning. Question-level annotations further identify the evidence target, counting structure, and required reasoning operation. The best evaluated VLM reaches 59.8% accuracy, compared with 98.2% for human annotators, and no evaluated open-weight model exceeds 45%. Stratified analyses reveal that failures are not uniformly distributed: structured target construction and action-grounded compositional reasoning form recurring bottlenecks across models. Zero-shot chain-of-thought prompting is not a reliable remedy: it recovers some errors but breaks previously correct answers, with effects that vary across models and task structures. VidNum therefore supports diagnostic analysis beyond a single aggregate score.