How Well Do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
Large language models (LLMs) are increasingly used in educational contexts, yet their ability to interpret authentic instructional discourse out-of-the-box remains unclear. We benchmark six state-of-the-art LLMs on classifying instructional moves in K-12 mathematics classroom transcripts annotated by expert educators (κ>0.90). We evaluated four prompting strategies, including zero-shot, one-shot, and few-shot prompts derived from the human coding manual. Zero-shot prompting achieved fair-to-moderate agreement (κ = 0.38–0.48, F1 = 0.45–0.53). Providing comprehensive examples improved performance for some models (e.g., κ = 0.48 to 0.58 for Claude 4.5 Opus; κ = 0.38 to 0.57 for Gemini 2.5 Pro), but gains were uneven and precision remained limited (best precision = 0.56, recall = 0.75). Errors were concentrated in constructs that require inference about instructor intent; for example, models confused Press for Reasoning with Press for Accuracy (42%–53% false-positive rates). Overall, our analysis found that LLMs demonstrate meaningful but limited capacity to identify aspects of instructional discourse, providing a baseline for educational discourse benchmarking and for designing more reliable annotation workflows. This work also points to a potential weakness in LLMs' ability to interpret key nuances of educational instruction.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex