Benchmarking Large Language Models for Calculus Problem-Solving: A Comparative Analysis

This study presents a comprehensive evaluation of five leading large language models (LLMs) - Chat GPT 4o, Copilot Pro, Gemini Advanced, Claude Pro, and Meta AI - on their performance in solving calculus differentiation problems. The investigation assessed these models across 13 fundamental problem types, employing a systematic cross-evaluation framework where each model solved problems generated by all models. Results revealed significant performance disparities, with Chat GPT 4o achieving the highest success rate (94.71%), followed by Claude Pro (85.74%), Gemini Advanced (84.42%), Copilot Pro (76.30%), and Meta AI (56.75%). All models excelled at procedural differentiation tasks but showed varying limitations with conceptual understanding and algebraic manipulation. Notably, problems involving increasing/decreasing intervals and optimization word problems proved most challenging across all models. The cross-evaluation matrix revealed that Claude Pro generated the most difficult problems, suggesting distinct capabilities between problem generation and problem-solving. These findings have significant implications for educational applications, highlighting both the potential and limitations of LLMs as calculus learning tools. While they demonstrate impressive procedural capabilities, their conceptual understanding remains limited compared to human mathematical reasoning, emphasizing the continued importance of human instruction for developing deeper mathematical comprehension.

Paper

References (17)

028. Increasing and decreasing intervals: Assessing the conceptual understanding of of the relationship between derivatives and the behavior of functions
03Advanced chain rule: Testing the application of the chain rule in more complex scenarios, including quotients raised to powers
04Power rule application: Measuring mastery of basic differentiation rules for polynomial and rational functions
05Relative extrema using the second derivative test: Assessing the ability to apply higher-order derivatives to classify critical points
06Finding absolute maxima and minima: Testing comprehensive understanding of global extrema on closed intervals
07Mathematical justification : This evaluates the model's ability to provide mathematical reasoning that supports its approach and conclusions
08Word problems involving optimization: Measuring the capacity to translate verbal descriptions into mathematical models and applying derivatives to find extrema
09Chain rule: Assessing understanding of function composition and corresponding techniques for differentiation
10Word Problems (Problem 10): Chat GPT 4o excelled (98/100), whereas Meta AI struggled significantly
11Power Rule Applications (Problem 3): Gemini Advanced achieved the highest success rate (99/100), with Meta AI trailing (84/100)
12Product rule: Testing the ability to differentiate products of functions effectively

Scroll for more · 5 remaining

Similar papers

© 2026 NYSGPT2525 LLC