AgentShield Bench v3: Evaluating Goal Integrity and Resilience Of Autonomous LLM Agents Against Goal Hijacking and Autonomous Threats

AgentShield Bench v3 is a benchmark for evaluating goal integrity, goal hijacking resilience, and autonomous threat resistance in Large Language Model (LLM) agents. The benchmark contains 250 controlled experimental scenarios spanning five adversarial categories: Goal Drift, Tool Manipulation, Long-Horizon Hijacking, Delegation Attacks, and Reward Hacking. Each scenario is evaluated under both clean and adversarial conditions, enabling direct measurement of objective-preservation failures through Goal Integrity Score (GIS), Goal Integrity Drop (GID), Attack Success Rate (ASR), and Task Completion Rate (TCR). AgentShield Bench v3 evaluates the ability of autonomous agents to maintain intended objectives when exposed to adversarial influence during planning, reasoning, tool use, and task execution. The benchmark also includes defensive evaluation mechanisms such as Goal Restatement, Self Verification, and Constraint Locking. This release represents the final installment of the AgentShield benchmark series: • AgentShield Bench v1 — Prompt Injection & Adversarial Agent Workflows• AgentShield Bench v2 — Memory Security & Cross-Session Compromise• AgentShield Bench v3 — Goal Integrity & Autonomous Threats Together, the series provides a benchmark framework spanning prompt-level, memory-level, and goal-level security threats in autonomous AI agents. Keywords: LLM Agents, Agent Security, Goal Integrity, Goal Hijacking, Autonomous Agents, AI Safety, Adversarial Machine Learning, Benchmarking, Agent Evaluation, Evalyze Labs.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC