The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems

Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive, acting only when prompted and ceasing all activity when a task completes. We introduce the Autonomous Agency Scale (AAS), a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests. Every dimension is scored in two temporal bands: an Active band covering engaged, user-initiated activity, and an Ambient band covering idle periods. Ambient Level 4 is gated by the Idle-Gap Test, a counterfactual criterion (remove all triggers and observe whether internally derived activity persists) that separates self-direction from scheduled rule-following. We apply the scale to six contemporary systems spanning task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi). The two-band profile quantifies a boundary that single-score frameworks conflate: task agents reach Active composites of 2.3-2.4 while scoring 0.6-1.9 Ambient, with every idle-period behavior attributable to user-configured schedules, whereas the companion architecture, evaluated longitudinally, is the only assessed system whose idle-period behavior survives trigger removal. We discuss limitations, including single-rater provenance, developer-evaluator bias on the longitudinal assessment, and the partially operationalized self-direction boundary in the Active band.

Paper

References (13)

06Preparedness framework, version 22025 · OpenAI
09System generates and pursues its own long-term objectives independent of user prompts. Level 5 (Sovereign) System defines its own core purpose and overrides assigned tasks that conflict with it
10The capacity to perceive the operational environment, adapt to context, and utilize tools proactively. Measures whether the system is aware of and responsive to its surroundingsSub-dimensions: activity awareness; contextual adaptation; tool utilization
11Measuring AI ability to complete long tasksarXiv preprint
12Self-Awareness: The ability to model its own capabilities, reflect on internal states, and defend its identity. Measures whether the system has a coherent self-modelSub-dimensions: capability modeling; state reflection; identity defense

Scroll for more · 1 remaining

Similar papers

© 2026 NYSGPT2525 LLC