vix.ing · top · new · best · stats

AI Testing Should Account for Sophisticated Strategic Behaviour

2025/08/19 by Vojtech Kovarik, Vojtěch Kovařík, Eric Olav Chen +9 · 1 voice · 2 citations
Computer Science · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Human-Automation Interaction and Safety #cs.AI #cs.GT

paper · pdf · doi:10.48550/arxiv.2508.14927

openalex publication_date 2025/08/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason strategically. Second, game-theoretic analysis can inform evaluation design by formalising and scrutinising the reasoning in evaluation-based safety cases. Drawing on examples from existing AI systems, a review of relevant research, and formal strategic analysis of a stylised evaluation scenario, we present evidence for these claims and motivate several research directions.

Citations

Cited by

Discussions

Related