Evaluating Large Language Models in Undergraduate Mathematics: Balancing Potentials and Pitfalls

Large language models (LLMs) are becoming integral to undergraduate student learning. Science, technology, engineering, and mathematics (STEM) majors are consulting LLMs to solve complex mathematical problems. Despite the quick, real-time feedback that LLMs can provide these students, LLMs can just as easily produce incorrect solutions that confuse students as they advance throughout their STEM degree programs. In this article, we analyze the performance of four LLMs—ChatGPT, Gemini Normal, Mistral, and Claude—at two different time points as they solve Calculus I, Calculus II, Linear Algebra, and Differential Equations questions taken from final exams administered in Fall 2022 at a large public university in the Mid-Atlantic region of the United States. Using a Solution Ranking Scale developed by the research team to evaluate the quality of the solutions that each LLM provides, we found the LLMs tended to find conceptual-based questions, with multiple symbols, more difficult to solve than application-based questions, which describe a mathematical context. These findings suggest that both LLM developers and higher education leaders consider the impact of variable quality in LLM mathematical problem-solving performance as more students continue to use LLMs as mathematics tutoring resources.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC