Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, a skill often needs to express more than task guidance: goals, input boundaries, permissions, human approval points, evidence requirements, output contracts, quality criteria, verification steps, and handoff rules. This paper proposes contractual skills, a GovernSpec-inspired design framework for organizing SKILL.md files as readable task contracts while preserving lightweight skill discovery and progressive loading. The framework clarifies the boundary between contractual skills, GovernSpec YAML contracts, Model Context Protocol (MCP) surfaces, tool adapters, runtime guardrails, tracing, and evaluation systems. We evaluate the framework with three offline empirical studies. The first text-generation experiment covers three enterprise skills, fifteen synthetic tasks, four instruction conditions, and eight generation models, producing 960 outputs and 1680 cross-judge score records. The second study is a public-skill A/B expansion: eight public skills are compared with contractual rewrites across forty-eight synthetic tasks, six generation models, two repeats, 1152 outputs, and two complete judge files. In this setting, contractual skills raise mean quality from 4.692 to 4.914 and reduce critical-error rate from 0.083 to 0.013. The third study is an offline tool-calling challenge with eight models and 192 simulated tool-call records. The results suggest that contractual skills are best understood as a governance layer that makes task intent, boundaries, and acceptance criteria explicit, not as a standalone safety mechanism.

Paper

References (19)

05Governspec: Runtime-independent contract compilation and offline acceptance testing for ai agent artifacts, 2026. URL https://ssrn.com/abstract=66748992026 · SSRN
06Support evaluation. Fields should map to checks such as required sections, forbidden commitments, privacy leakage, uncertainty marking, and handoff completeness4.2 Field model Table 1 summarizes
07It reports a market-validated skill A/B expansion with eight public skills, 1152 outputs, and two complete judge files
08Towards secure agent skills:Architecture, threat taxonomy, and security analysis
09It clarifies the boundary between contractual skills, GovernSpec YAML, MCP, tool adapters, runtime guardrails, tracing, and evaluation
10Support gradual adoption. Teams can start by rewriting existing skills, then add GovernSpec YAML and offline tests for high-risk workflows
11Anthropic:/
12Separate skill instructions from runtime enforcementSkills express intent; adapters, guardrails, and permissions enforce behavior

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC