Empirical Evaluation of Deep Learning Models for Knowledge Tracing: Of Hyperparameters and Metrics on Performance and Replicability
New knowledge tracing models are continuously being proposed, even at a pace where state-of-theart<br> models cannot be compared with each other at the time of publication. This leads to a situation<br> where ranking models is hard, and the underlying reasons of the models’ performance – be it architectural<br> choices, hyperparameter tuning, performance metrics, or data – is often underexplored. In this<br> work, we review and evaluate a body of deep learning knowledge tracing (DLKT) models with openly<br> available and widely-used data sets, and with a novel data set of students learning to program. The<br> evaluated knowledge tracing models include Vanilla-DKT, two Long Short-Term Memory Deep Knowledge<br> Tracing (LSTM-DKT) variants, two Dynamic Key-Value Memory Network (DKVMN) variants,<br> and Self-Attentive Knowledge Tracing (SAKT). As baselines, we evaluate simple non-learning models,<br> logistic regression and Bayesian Knowledge Tracing (BKT). To evaluate how different aspects of DLKT<br> models influence model performance, we test input and output layer variations found in the compared<br> models that are independent of the main architectures. We study maximum attempt count options, including<br> filtering out long attempt sequences, that have been implicitly and explicitly used in prior studies.<br> We contrast the observed performance variations against variations from non-model properties such as<br> randomness and hardware. Performance of models is assessed using multiple metrics, whereby we also<br> contrast the impact of the choice of metric on model performance. The key contributions of this work are<br> the following: Evidence that DLKT models generally outperform more traditional models, but not necessarily<br> by much and not always; Evidence that even simple baselines with little to no predictive value<br> may outperform DLKT models, especially in terms of accuracy – highlighting importance of selecting<br> proper baselines for comparison; Disambiguation of properties that lead to better performance in DLKT<br> models including metric choice, input and output layer variations, common hyperparameters, random<br> seeding and hardware; Discussion of issues in replicability when evaluating DLKT models, including<br> discrepancies in prior reported results and methodology. Model implementations, evaluation code, and<br> data are published as a part of this work.
Paper
References (100)
Scroll for more · 38 remaining