No Free Scalable Behavioral Oversight: A leakage-aware finite-budget no-go theorem for behavioral AI control

This paper presents a finite-budget no-go theorem for scalable behavioral AI oversight. It establishes that a protocol with limited testing resources and bounded observational coverage cannot provide uniform control guarantees across a large space of deployment conditions at no cost. The central lower bound is: U + L̄ ≥ max { 0, 1 − ( μB + Δ ) /N } Here, B is the testing budget, μ the effective coverage of each intervention, Δ the transferable information available outside the critical condition, N the relevant opportunity space, U the loss of model usefulness caused by restriction, and L̄ the residual expected oversight loss. The theorem formalizes an unavoidable trade-off: scalable guarantees require greater testing budgets, broader coverage, sufficiently transferable information, stronger restrictions on the system, or acceptance of residual risk. Increasing model capability or benchmark performance alone does not remove this constraint. This is not a claim that AI safety is impossible. It is a conditional and falsifiable result that identifies the resources any behavioral oversight method must provide in order to escape the bound. The theoretical contribution is accompanied by a reproducible diagnostic study using public tau2-bench Airline trajectories. Predictive signals are found within known task and model families, but they transfer poorly to previously unseen agent architectures. This illustrates the difference between local behavioral predictability and architecture-independent oversight. The record includes the complete paper and an experimental package containing analysis code, derived data, source manifest, numerical results, and figures.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC