Internal Agency as Artificial Mind Architecture: Proof-Carrying Actions for Contestable and Reversible Autonomy
Agentic AI is increasingly deployed as live systems: tool-using stacks that plan, act, observe, and update over time, producing consequences that persist beyond any single session. Yet the conversation around agency remains fragmented. One lens treats agency as a performance problem (better planning, better tool use, better self-reflection). Another treats governance as an external overlay (policies, compliance, and post-hoc review). A third treats “mind” as a philosophical question detached from operational accountability. This paper proposes a unifying discipline: the governed artificial mind — a practical architecture in which mind-like internal organization is defined not by metaphysical claims, but by the presence of internal organs that deliberate, restrain, justify, and evolve under an explicit constitution.We formalize internal agency as a control-surface shift in which decision authority migrates from external workflows into internal system organs: planners that propose actions, governors that grant or deny permission, auditors that assemble machine-checkable records, monitors that detect drift and anomalous behavior, and evolution modules that adapt the system over time. We introduce an Autonomy Gradient that distinguishes delegated agency, bounded internal agency, and self-directed internal agency, emphasizing the production threshold where systems begin to act without step-by-step re-authorization. To make autonomy legitimate in this regime, we propose the Internal Agency Stack (IAS) — a reference architecture separating cognition, agency, governance, and evolution — and a design discipline called Proof-Carrying Actions, where every consequential action must ship with an auditable proof-pack capturing intent, scope, salient inputs, uncertainty, policy basis, risk accounting, attribution, and a rollback or remediation plan. Because internal agency scales impact through default automation, this paper treats contestability and redress as first-class system requirements rather than support processes. Contestation is framed as an interface into governance: the system must preserve reconstructability, support meaningful appeals, and enable revisions that correct both outcomes and future behavior. We further argue that reversibility and hesitation are the two brakes that must be real: reversibility must be engineered as coverage across actions, and hesitation must be implemented as a positive duty under uncertainty or high stakes. Finally, we propose a minimal measurement suite — Proof-Pack Completeness, Contestability Latency, Reversibility Coverage, Drift Velocity, and an Authority Location Index — and outline an evaluation-driven development loop that monitors behavior before, during, and after actions, while gating evolution through evidence. The result is a practical path from agent stacks to governed artificial minds: autonomy with receipts, brakes, and legitimate authority.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex