Understanding the role of large language models in supporting mathematics education

Large Language Models (LLMs) are redefining mathematics education, including not only student-facing tutors but also educator-centered tools that automate instructional tasks and support professional development. This systematic review synthesizes findings from 131 empirical and technical studies published between January 2022 and February 2025. Guided by PRISMA protocols and BERTopic modeling, we investigated three key questions: (1) How do LLMs perform on mathematical tasks across varying levels of difficulty? (2) In what ways are LLMs being used to support mathematics educators? (3) What are the limitations and concerns of applying LLMs in mathematics education? BERTopic identified four dominant themes: (a) evaluation of LLM mathematical reasoning, (b) LLM-supported mathematics education, (c) automation of mathematics educational tasks, and (d) LLM capability in mathematical word problems. Across benchmarks (e.g., GSM8K, SVAMP, MATH), LLM performance is strong on many primary and secondary tasks, especially when combined with remediation approaches such as prompt engineering, fine-tuning, and tool/code-based verification, but reliability decreases as problems require abstract reasoning, multi-step planning, or domain-specific formalisms (notably calculus, statistics/probability, geometry/spatial reasoning, and some non-English settings). Synthesizing findings through an Intelligent Tutoring Systems (ITS) lens, we argue that LLMs are most defensible as generative components within systems that provide instructional control, verification, and student-model inference, rather than as autonomous authorities for mathematical correctness or pedagogy. Key risks include hallucination, limited pedagogical sensitivity, uneven multilingual robustness, and over-reliance that may weaken learners’ independent reasoning. We conclude with implications for evaluation and adoption, emphasizing verifiability, process-sensitive assessment, and educator-centered governance.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC