Large language models (LLMs) are increasingly deployed in academic and professional contexts where epistemic restraint, boundary recognition, and instruction adherence are as important as raw capability. However, existing evaluations primarily emphasize task performance and fluency, offering limited insight into how models behave under explicit epistemic constraints. This paper presents a pilot methodological study applying controlled experimental design principles to LLM instruction-following behavior. Two role-conditioned instruction profiles-Academic and Devil's Advocate-were applied to a single model under identical prompts and clean-slate conditions. While both profiles frequently complied with surface-level constraints, they diverged in approach or content in most tests and exhibited outcome-level divergence in three of eight tests. Critical failure modes were observed, including speculative completion and critique overflow. Overall compliance rates were 5/8 for the Academic profile and 4/8 for the Devil's Advocate profile, suggesting that instruction adherence is achievable but not reliable. The contribution is methodological rather than empirical: a reproducible, negative-result-tolerant framework for diagnosing epistemic behavior and instruction hierarchy failures in LLMs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex