Agentic AI Cyber Subversion: The Semantic Layer Integrity Attack as a New Threat Class Against the Reasoning Layer
Agentic AI introduces a cyber risk that does not begin with compromised credentials, poisoned training data, prompt injection, unauthorized tool use, or anomalous output. It begins when an autonomous system resolves the meaning of a regulated term across enterprise systems before execution, and that resolved meaning diverges from the definition the institution authorized. The corrupted object is not the input, the output, the model, the tool call, or the log. It is the resolved operational meaning. This paper defines Agentic Workflow Subversion as the enterprise risk surface created when reasoning-layer drift propagates across workflows, systems, and control boundaries.It defines the Semantic Layer Integrity Attack (SLIA) as the deliberate adversarial form of that risk: a cross-system integrity attack in which an actor manipulates the semantic conditions under which an agent resolves authorization, eligibility, clearance, risk, or control status, while every contributing system continues to behave correctly. It is a failure of control integrity without system compromise. The attack is cyber-relevant because it produces a clean control record. The network is not breached. The model is not necessarily altered. The prompt may be benign. The tool call may be authorized. The output may be well formed. The audit trail may be complete. Yet execution proceeds under a corrupted operational interpretation that no human or institution authorized. It situates the attack against existing agent security controls, including identity, workload authentication, prompt and input defenses, tool governance, memory and retrieval controls, output filtering, auditability, and defense-in-depth architectures. It shows that these controls are necessary but incomplete, because they govern the artifacts around agentic reasoning rather than the meaning resolved by the reasoning layer itself. It then maps the gap against current cyber and AI governance taxonomies, including NIST adversarial machine learning, the OWASP agentic AI security corpus, MITRE ATLAS, STRIDE, the Five Eyes agentic AI security guidance, and financial sector supervisory expectations, in the disciplined posture that each framework is authoritative for the scope it declares and the reasoning layer is the adjacent surface it does not directly observe. The contribution is a cyber threat class definition and a control surface argument. Agentic Workflow Subversion names the risk surface. The Semantic Layer Integrity Attack names its adversarial exploitation path. Semantic Substrate Validation and the Semantic Control Plane (SCP), operating within the broader Agentic Governance Model (AGM), name the runtime governance surface required to detect, constrain, and evidence semantic integrity before execution proceeds. It closes with detection and test conditions, control tailoring, and supervisory implications, so that the threat class is operationally actionable for a cybersecurity function rather than only conceptually defined. This work is also available on SSRN as Working Paper No. 6926219. https://ssrn.com/abstract=6926219
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex