Large language models (LLMs) are increasingly being used in education, yet their correctness alone does not capture the quality, reliability, or pedagogical validity of their problem-solving behavior, especially in mathematics, where multi-step logic, symbolic reasoning, and conceptual clarity are critical. Conventional evaluation methods largely focus on final answer accuracy and overlook the reasoning process. To address this gap, we introduce a novel interpretability framework for analyzing LLM-generated solutions using undergraduate calculus problems as a representative domain. Our approach combines reasoning flow extraction and decomposing solutions into semantically labeled operations and concepts with prompt ablation analysis to assess input salience and output stability. Using structured metrics such as reasoning complexity, phrase sensitivity, and robustness, we evaluated the model behavior on real Calculus I–III university exams and compared it with the performances of students enrolled in the courses. Our findings revealed that LLMs often produce syntactically fluent yet conceptually flawed solutions with reasoning patterns sensitive to prompt phrasing and input variation. This framework enables a fine-grained diagnosis of reasoning failures, supports curriculum alignment, and informs the design of interpretable AI-assisted feedback tools. The framework was evaluated on Gemma 3, an open-access large language model, across zero-shot, retrieval-augmented generation, and contextual retrieval configurations, using nine real undergraduate calculus examinations from three course levels. To our knowledge, this is the first paper to apply a combined reasoning flow decomposition and prompt ablation framework to real undergraduate calculus examinations, benchmarked against actual student cohort performance, laying the foundation for the transparent and responsible deployment of AI in STEM learning environments.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex