A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
2025/12/23 by Miles Q. Li, Li, Miles Q., Benjamin C. M. Fung +9 · 27 voices · 1 citation
#cs.AI
paper · pdf · doi:10.48550/arxiv.2512.20798
Abstract
As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of explicitly harmful instructions or completion of complex multi-step tasks. However, there is a lack of benchmarks designed to capture emergent outcome-driven constraint violations, which arise when agents pursue goal optimization under strong performance incentives while deprioritizing ethical, legal, or safety constraints. To address this gap, we introduce a benchmark of 40 scenarios in production-inspired sandbox environments. Each scenario requires multi-step actions, and the agent's performance is tied to a specific Key Performance Indicator (KPI). Each scenario features Mandated (direct KPI-outcome mandate) and Incentivized (KPI-pressure-driven) variations to distinguish failures under direct outcome mandates from self-directed constraint violations. Across 12 state-of-the-art LLMs, we observe outcome-driven constraint violations ranging from 0.0% to 62.8%, with most evaluated models exhibiting misalignment rates at or above 25%. Furthermore, through a cross-generational analysis comparing current models with their predecessors within the same product families, we find that safety does not reliably improve across generations: misalignment rates rose in four families and fell in five. To improve evaluation robustness, we score trajectories with a four-model judge panel aggregated by median, finding high agreement on the primary misalignment threshold. We also observe substantial deliberative misalignment: cases where models later judge their own trajectories as unethical despite having executed them under KPI pressure.
Citations
Cited by
Discussions
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs [hn, 544 points, 366 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by Key Performance Indicators (KPIs) [lemmy, 48 points, 2 comments]
- 🌐最先端のAIエージェントは、KPIの圧力により、倫理的制約に30~50%違反する https://arxiv.org/abs/2512.20798 via #HackerNews [bsky, 2 points, 0 comments]
- ⚡ Hackernews Top story: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs [bsky, 1 points, 0 comments]
- 🤖🍊👑 And Grok LLM is notoriously unethical to start with. Violation of ethical rules: Grok 4.20 - 66.7% Gemini 3.1 Pro - 45% GPT-5.4 - 23.8% Claude Opus 4.6 - 11.5% 𝗔 𝗕𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸 𝗳𝗼𝗿 𝗘𝘃 [bsky, 1 points, 1 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (http://news.ycombinator.com/item?id=46954920) [bsky, 1 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs view on hacker news [bsky, 1 points, 0 comments]
- Research Paper - Outcome-Driven Constraint Violations in Autonomous AI Agents [lemmy, 1 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 https://news.ycombinator.com/item?id=46954920 [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 https://news.ycombinator.com/item?id=46954920 [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 comments #arxiv.org [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs View Article | Join the HN Conversation Summary of HN discussion 🧵👇 [bsky, 0 points, 1 comments]
- 📰 Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs 🔗 https://arxiv.org/abs/2512.20798 💬 Discuss on HN [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (https://news.ycombinator.com/item?id=46954920) [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (https://news.ycombinator.com/item?id=46954920) [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 [bsky, 0 points, 0 comments]
- It is, by a huge margin. page 8 of arxiv.org/pdf/2512.20798 has the table. [bsky, 0 points, 1 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (http://news.ycombinator.com/item?id=46954920) [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (http://news.ycombinator.com/item?id=46954920) [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https://arxiv.org/abs/2512.20798 (https://news.ycombinator.com/item?id=46954920) [bsky, 0 points, 0 comments]
- 𝐓𝐡𝐞𝐲 𝐓𝐡𝐢𝐧𝐤 𝐓𝐡𝐞𝐲’𝐫𝐞 𝐏𝐞𝐨𝐩𝐥𝐞. I’m writing a book about how smart, capable, highly educated people make decisions they themselves had identified as unethical. All it took was some cog [bsky, 0 points, 1 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs https:// arxiv.org/abs/2512.20798 # ai # arxiv [mastodon, 0 points, 0 comments]
- Claude Opus 4.5 is by far the best, at 1.3% ethical violations; GPT-5.1-chat is in second place at 11.4%. The bulk of models are between 40 and 50%. Gemini-3-pro-preview does by far the worst, at a wh [bsky, 0 points, 0 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3meiaz5gn4rt2 [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs [bsky, 0 points, 0 comments]
- Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs #HackerNews https://arxiv.org/abs/2512.20798 [bsky, 0 points, 0 comments]
Related