This paper examines whether current artificial intelligence alignment methods can scale in step with rapidly increasing model capability. Using a general framework for regulatory burden scaling in complex systems, the study analyzes alignment as a form of internal regulation whose cost grows with system activity. The core result is a conditional scaling theorem: if alignment effort grows sublinearly with training compute while capability grows superlinearly, then a finite computational threshold must exist beyond which internal safety margins collapse. The paper formalizes this threshold using power-law scaling relationships and derives explicit expressions linking capability growth, alignment overhead, and failure risk. Drawing on published scaling trends and benchmark data, the work argues that prevailing alignment techniques—such as reinforcement learning from human feedback, red-teaming, and post-hoc safety filters—are architecturally predisposed to sublinear scaling. Under continued compute growth, this mismatch produces what the paper terms an “alignment desert”: a regime in which increased capability systematically outpaces the system’s ability to remain aligned. The paper does not claim empirical proof of catastrophic failure, but instead presents falsifiable predictions, quantitative diagnostics, and clear conditions under which current methods would fail. It also proposes a class of architectures—Regulation-Embedded Transformers—as a potential pathway toward superlinear regulatory scaling, outlining how alignment mechanisms can be integrated directly into model computation rather than appended post-hoc. Overall, the work frames AI alignment not solely as a training or oversight problem, but as an architectural scaling constraint. It provides a mathematically grounded lens for evaluating alignment strategies, identifying risk thresholds, and designing systems whose safety properties improve with scale rather than degrade.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex