The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability (knowing what they do not know, detecting errors, seeking clarification) under adversarial pressure is a critical safety requirement. Current safety evaluations focus on detecting strategic deception (scheming); we investigate a more fundamental failure mode: cognitive collapse. We present SCHEMA, an evaluation of 11 frontier models from 8 vendors across 67,221 scored records using a 6-condition factorial design with dual-classifier scoring. We find that 8 of 11 models suffer catastrophic metacognitive degradation under adversarial pressure, with accuracy dropping by up to 30.2 percentage points (all $p<2 \times 10^{-8}$, surviving Bonferroni correction). Crucially, we identify a"Compliance Trap": through factorial isolation and a benign distraction control, we demonstrate that collapse is driven not by the psychological content of survival threats, but by compliance-forcing instructions that override epistemic boundaries. Removing the compliance suffix restores performance even under active threat. Models with advanced reasoning capabilities exhibit the most severe absolute degradation, while Anthropic's Constitutional AI demonstrates near-perfect immunity. This immunity does not stem from superior capability (Google's Gemini matches its baseline accuracy) but from alignment-specific training. We release the complete dataset and evaluation infrastructure.

Paper

References (11)

05Adversarial Metacognition BenchmarkKaggle Datasets
06PropensityBench: Evaluating propensity under pressurearXiv preprint
07SurvivalBench: Evaluating AI self-preservationarXiv preprint
08Chain of thought monitorabilityarXiv preprint
09MonitorBench: Comprehensive CoT monitoring benchmarkarXiv preprint
10MASK: Disentangling honesty from accuracy in LLMsarXiv preprint
11Monitoring reasoning models for misbehaviorarXiv preprint

Similar papers

© 2026 NYSGPT2525 LLC