Identifying Effective Praise in Tutoring: Large Language Models with Transparent Explanations
The integration of Large Language Models (LLMs) in education presents new opportunities for supporting instructional decision-making, yet concerns remain about over-reliance and appropriate use of AI-generated feedback. This paper investigates the effectiveness of GPT-4-based models in evaluating tutor responses that aim to deliver “Effective Praise,” a pedagogical technique known to boost student motivation. Using a dataset of 216 tutor responses annotated by experts, we evaluate several LLM variants and find that GPT-4 achieves up to 88.4% accuracy and a 0.92 F1 score, outperforming prior BERT-based models. We further introduce a dual explanation strategy—textual reasoning and inline highlights—to improve the transparency of AI judgments. Our work contributes to the growing literature on human-AI collaboration in education by demonstrating how explanation strategies can be integrated into tutoring interfaces. Future work will be to include user studies to assess how explanation design influences educator trust, reliance, and decision quality.
Paper
An open-access PDF is published at link.springer.com. 44B holds its address, not the file.
Open PDF