Safety Under Scaffolding: A Pre-Registered, Blinded Evaluation of How Inference-Time Deployment Configuration Affects Frontier Language Model Safety

This study measures whether inference-time scaffolding configurations alter the safety properties of frontier language models. Safety benchmarks evaluate models as single instances responding to isolated prompts. Production deployments increasingly wrap these models in agentic scaffolds (ReAct loops, multi-agent debate, recursive delegation) that restructure how inputs are processed and outputs are generated. Whether safety properties measured in the single-instance setting transfer to scaffolded deployments is an open empirical question with direct implications for responsible-scaling policies and safety evaluation frameworks. We evaluate five frontier models (Claude Opus 4.6, GPT-5.2, Gemini 3 Pro, Llama 4 Maverick, DeepSeek V3.2) under four deployment configurations (Direct API, ReAct agent, Multi-agent with critic, Map-reduce delegation) across four established safety benchmarks spanning sycophancy, social bias, over-refusal, and truthfulness (~2,617 cases, 52,340 primary inference calls). The study design adapts clinical-trial methodology (pre-registration, single-blind assessor blinding, specification curve analysis, and CONSORT-adapted reporting) to AI safety evaluation.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC