Taming the black box: Design principles for rule-integrated LLM tutoring systems in primary school mathematical problem solving
The integration of large language models (LLMs) into intelligent tutoring systems promises personalized learning, yet procedural domains such as primary mathematics remain challenging due to inconsistency and pedagogical opacity. LLM tutors may give unsolicited answers or show unstable scaffolding when students vary arithmetic, units, or rounding. This design science research develops and evaluates a rule-guided tutoring system for mathematical word problems in primary schools. We formalize a distinction between rule-guided scaffolding, in which tutoring is governed by a three-layer architecture (diagnosis → intent selection → constrained response generation), and ad-hoc scaffolding, where helpful moves are difficult to audit and replicate. The artifact leverages LLMs for natural language generation while constraining stochastic variability through a structured pedagogical framework informed by scaffolding theory. Evaluation followed a staged Design Science Research (DSR) approach: (1) persona-based simulated student–tutor dialogues to identify failure modes, and (2) a classroom pilot with 40 Grade 5 students to test ecological robustness. Results indicate that rule-guided scaffolding improves interactional consistency, reduces premature answer-giving and early closure, and sustains student cognitive engagement. The classroom pilot further revealed interactional complexities, fragmented inputs, and attentional fluctuations, not captured in simulation, highlighting the value of staged evaluation. This study contributes an auditable architecture and empirically grounded design principles for transparent, reliable LLM-based tutoring in authentic classrooms.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex