Automated Assessment of Handwritten Math Problems: A Comparison of Prompting Strategies for Open and Closed-source LLMs
Assessing handwritten mathematical solutions is essential for identifying students’ weaknesses and fostering personalized learning. However, scaling such assessment remains challenging for Learning Analytics, which has traditionally focused on digital or typed data. The current study investigated the potential of Large Language Models (LLMs) to automating the assessment of handwritten mathematical solutions and explore how they can be incorporated into large scale learning analytics pipelines. We curated 300 student solution images, annotated them using to a taxonomy of math error types, and compared open-source (Qwen2.5-7B and Gemma3 12B-IT) and closed-source (Gemini 2.0 Flash and GPT-4) LLMs. Two prompting strategies were tested: from adapted from related work and one tailed to the taxonomy of math error types using established prompt design principles. The results revealed that LLMs, particularly Gemini, achieved strong to moderate performance in diagnosing and classifying student errors, while exposing recurring model specific errors. These findings highlight both the promise and limitations of LLMs for integrating handwritten work in LA and recommend that learning analytics practitioners and researchers combine careful model selection, principled prompt design, and error-level analysis to develop AI-powered LA systems that are accurate, equitable, and pedagogically actionable.
Paper
Full text
Automated Assessment of Handwritten Math Problems: A Comparison of Prompting Strategies for Open and Closed-source LLMs
OpenAlex · Intelligent Tutoring Systems and Adaptive Learning · 2026
Abstract
Assessing handwritten mathematical solutions is essential for identifying students’ weaknesses and fostering personalized learning. However, scaling such assessment remains challenging for Learning Analytics, which has traditionally focused on digital or typed data. The current study investigated the potential of Large Language Models (LLMs) to automating the assessment of handwritten mathematical solutions and explore how they can be incorporated into large scale learning analytics pipelines. We curated 300 student solution images, annotated them using to a taxonomy of math error types, and compared open-source (Qwen2.5-7B and Gemma3 12B-IT) and closed-source (Gemini 2.0 Flash and GPT-4) LLMs. Two prompting strategies were tested: from adapted from related work and one tailed to the taxonomy of math error types using established prompt design principles. The results revealed that LLMs, particularly Gemini, achieved strong to moderate performance in diagnosing and classifying student errors, while exposing recurring model specific errors. These findings highlight both the promise and limitations of LLMs for integrating handwritten work in LA and recommend that learning analytics practitioners and researchers combine careful model selection, principled prompt design, and error-level analysis to develop AI-powered LA systems that are accurate, equitable, and pedagogically actionable.
References (26)
Scroll for more · 14 remaining