VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI
Educational AI systems increasingly rely on large language models to support student writing and inquiry, yet enforcing safety and instructional constraints during open-ended, multi-turn interaction remains challenging. Existing approaches commonly embed such constraints within conversational prompts or rely on static filtering. Over time, these approaches may become sensitive to user interaction, making it difficult to monitor and audit when students are able to circumvent or otherwise attempt to violate such measures. We introduce VETTING, a dual-LLM framework that separates response generation from policy verification and applies explicit policy checks at runtime. The framework is illustrated through a grounded instantiation that enforces instructional and safety constraints without exposing policy specifications during interaction. We evaluate VETTING through an in situ deployment in a middle school writing activity, combining analysis of student–AI interaction behavior, human audit of verification outcomes, and characterization of computational overhead. In this deployment, VETTING achieved a precision of.943, recall of.913, and F1 score of.928, and was associated with an estimated 91.2% reduction in inappropriate content exposure with a 19.6% increase in token usage (628,011 tokens). These results suggest that policy-isolated runtime verification can support the analysis and management of educational AI behavior under authentic classroom use.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex