Persistent Performance Models for Failure Regime Analysis in Simulated LLM Agents: A Controlled Ablation Study
We present a simulation study of evaluation-induced structure in LLM agent performance measurement. PPM-EA (Persistent Performance Model — Execution Anchored Drift Analysis) is a controlled framework in which monitoring intensity, verifier reliability, and prompt entropy are varied as independent experimental factors; all findings are properties of the simulation's parameterization. Across two complementary experiments — a 2^3 factorial grid (N=160 task records, Experiment 1) and a five-condition verifier study (N=400, Experiment 2) — we examine how these factors interact to produce measurable performance collapse. Experiment 1 yields partial structure: verifier reliability is the dominant single factor (eta^2=0.395), the three-way interaction explains eta^2=0.472 of collapse variance, yet cross-seed transition stability (0.706) and invariance (0.540) fall short of strong-structure thresholds. Experiment 2 isolates the verifier to an observation-only channel and reveals a dominance inversion: monitoring intensity rises to eta^2=0.238 on verifier-agnostic (true) collapse while verifier type drops to eta^2=0.005; 35.7% of observed collapse variance is attributable to verifier observation artifact. A small external probe with a real local LLM (N=72 step records, gemma4:latest) provides directional confirmation: a format-biased verifier rejects 100% of correct low-monitoring responses while the ground-truth proxy passes 77.8%, yielding a format-artifact bias of +0.722 in a real inference loop. Code and data: https://github.com/leandroconsolaro66/ppm-ea
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex