vix.ing · top · new · best · stats · spec

One Human, N Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence

2026/07/30 by Cesare Zavattari, Alessandro Tommasi, Giuseppe Prencipe
Computer Science · #cs.AI

paper · pdf

arxiv created 2026/07/30 · arxiv updated 2026/07/31

Abstract

A single human must audit N LLM agents under a budget of B ≪ N audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold δ^* past which confidence-ranked auditing is worse than random. Two a-priori expectations reverse: δ^* rises as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for vacuous oversight, and replaying policies on recorded traces confirms the ordering.

Related