Automated Assessment of Handwritten Math Problems: A Comparison of Prompting Strategies for Open and Closed-source LLMs

Assessing handwritten mathematical solutions is essential for identifying students’ weaknesses and fostering personalized learning. However, scaling such assessment remains challenging for Learning Analytics, which has traditionally focused on digital or typed data. The current study investigated the potential of Large Language Models (LLMs) to automating the assessment of handwritten mathematical solutions and explore how they can be incorporated into large scale learning analytics pipelines. We curated 300 student solution images, annotated them using to a taxonomy of math error types, and compared open-source (Qwen2.5-7B and Gemma3 12B-IT) and closed-source (Gemini 2.0 Flash and GPT-4) LLMs. Two prompting strategies were tested: from adapted from related work and one tailed to the taxonomy of math error types using established prompt design principles. The results revealed that LLMs, particularly Gemini, achieved strong to moderate performance in diagnosing and classifying student errors, while exposing recurring model specific errors. These findings highlight both the promise and limitations of LLMs for integrating handwritten work in LA and recommend that learning analytics practitioners and researchers combine careful model selection, principled prompt design, and error-level analysis to develop AI-powered LA systems that are accurate, equitable, and pedagogically actionable.

Paper

Full text

PDF

Automated Assessment of Handwritten Math Problems: A Comparison of Prompting Strategies for Open and Closed-source LLMs

OpenAlex · Intelligent Tutoring Systems and Adaptive Learning · 2026

Abstract

Assessing handwritten mathematical solutions is essential for identifying students’ weaknesses and fostering personalized learning. However, scaling such assessment remains challenging for Learning Analytics, which has traditionally focused on digital or typed data. The current study investigated the potential of Large Language Models (LLMs) to automating the assessment of handwritten mathematical solutions and explore how they can be incorporated into large scale learning analytics pipelines. We curated 300 student solution images, annotated them using to a taxonomy of math error types, and compared open-source (Qwen2.5-7B and Gemma3 12B-IT) and closed-source (Gemini 2.0 Flash and GPT-4) LLMs. Two prompting strategies were tested: from adapted from related work and one tailed to the taxonomy of math error types using established prompt design principles. The results revealed that LLMs, particularly Gemini, achieved strong to moderate performance in diagnosing and classifying student errors, while exposing recurring model specific errors. These findings highlight both the promise and limitations of LLMs for integrating handwritten work in LA and recommend that learning analytics practitioners and researchers combine careful model selection, principled prompt design, and error-level analysis to develop AI-powered LA systems that are accurate, equitable, and pedagogically actionable.

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC