Why Large Language Models Alone Fall Short for Responsible Learner Modeling in K–12 Tutoring: A Case Study
The rapid rise of large language model (LLM)-based tutors in K–12 education has led to the misconception that generative models can replace traditional learner modelling and act as general-purpose engines for adaptive instruction. This is especially problematic in K–12 settings, which the EU AI Act classifies as a high-risk domain requiring responsible design. Motivated by concerns surrounding the role of learner modelling in responsible AI-powered tutoring, this study synthesises existing research evidence on key limitations of LLM-based tutoring systems and then presents an empirical case study investigating one critical aspect of these concerns: the accuracy, reliability, and temporal coherence of assessing learners’ evolving knowledge over time. To this end, we compare a deep knowledge tracing (DKT) model with a widely used LLM (with and without fine-tuning) that has demonstrated competitive performance in tutoring-related tasks, using a large-scale open-access dataset. Our findings show that DKT achieves the highest discrimination performance (AUC = 0.83) onnext-step correctness predictionand consistently outperforms the LLM across evaluation settings. Although fine-tuning improves the LLM’s AUC by about 8% over the zero-shot baseline, it still remains 6% below DKT and produces higher early-sequence errors, precisely where incorrect predictions would be most harmful for adaptive learner support. Temporal-coherence analyses further reveal that while DKT maintains stable, directionally correct mastery updates, LLM variants display substantial temporal weaknesses, including smooth but wrong-direction updates, and, even after fine-tuning, remain inconsistent and unable to match DKT’s temporal stability. We also illustrate that these issues persist even though fine-tuned LLMs required nearly 198 hours of continuous high-compute training, far exceeding the computational demands of the lightweight DKT model. Our qualitative analysis ofmulti-skill mastery estimationfurther shows that, even after fine-tuning, the LLM produced unstable and inconsistent mastery trajectories, whereas DKT maintained smooth and coherent multi-skill updates. Collectively, these findings suggest that LLMs alone are unlikely to achieve the same positive effect sizes observed in decades long intelligent tutoring systems literature. Rather than replacing learner modelling, LLMs may be more appropriately deployed as pedagogical interfaces or content generators paired with dedicated learner modelling components to ensure responsible, accurate, reliable, and pedagogically sound support.
Paper
Full text
Why Large Language Models Alone Fall Short for Responsible Learner Modeling in K–12 Tutoring: A Case Study
OpenAlex · Intelligent Tutoring Systems and Adaptive Learning · 2026
Abstract
The rapid rise of large language model (LLM)-based tutors in K–12 education has led to the misconception that generative models can replace traditional learner modelling and act as general-purpose engines for adaptive instruction. This is especially problematic in K–12 settings, which the EU AI Act classifies as a high-risk domain requiring responsible design. Motivated by concerns surrounding the role of learner modelling in responsible AI-powered tutoring, this study synthesises existing research evidence on key limitations of LLM-based tutoring systems and then presents an empirical case study investigating one critical aspect of these concerns: the accuracy, reliability, and temporal coherence of assessing learners’ evolving knowledge over time. To this end, we compare a deep knowledge tracing (DKT) model with a widely used LLM (with and without fine-tuning) that has demonstrated competitive performance in tutoring-related tasks, using a large-scale open-access dataset. Our findings show that DKT achieves the highest discrimination performance (AUC = 0.83) on next-step correctness prediction and consistently outperforms the LLM across evaluation settings. Although fine-tuning improves the LLM’s AUC by about 8% over the zero-shot baseline, it still remains 6% below DKT and produces higher early-sequence errors, precisely where incorrect predictions would be most harmful for adaptive learner support. Temporal-coherence analyses further reveal that while DKT maintains stable, directionally correct mastery updates, LLM variants display substantial temporal weaknesses, including smooth but wrong-direction updates, and, even after fine-tuning, remain inconsistent and unable to match DKT’s temporal stability. We also illustrate that these issues persist even though fine-tuned LLMs required nearly 198 hours of continuous high-compute training, far exceeding the computational demands of the lightweight DKT model. Our qualitative analysis of multi-skill mastery estimation further shows that, even after fine-tuning, the LLM produced unstable and inconsistent mastery trajectories, whereas DKT maintained smooth and coherent multi-skill updates. Collectively, these findings suggest that LLMs alone are unlikely to achieve the same positive effect sizes observed in decades long intelligent tutoring systems literature. Rather than replacing learner modelling, LLMs may be more appropriately deployed as pedagogical interfaces or content generators paired with dedicated learner modelling components to ensure responsible, accurate, reliable, and pedagogically sound support.