Large Language Models (LLMs) are increasingly being integrated into educational and problem-solving domains, prompting investigations into their mathematical reasoning capabilities. A comparative study of four prominent LLMs—GPT-2, T5, BART, and Tiny BERT—evaluated their performance on algebraic equation solving tasks of varying complexity. The assessment encompassed polynomial roots, one-dimensional equations, two-dimensional systems, and composed expressions, with models evaluated based on solution accuracy and consistency. Results demonstrated a positive correlation between model size and performance, with GPT-2 exhibiting exceptional proficiency (exceeding 97% accuracy) even on complex equations. BART and T5 similarly displayed strong capabilities, while Tiny BERT showed measurable but less substantial improvement. These findings suggest significant potential for LLMs as computational tools for mathematical reasoning and automated algebraic problem-solving applications.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex