The One-Shot Ceiling: Comparing RAG and Fine-Tuning Architectures for AI-Assisted Math Mentoring

Providing individualized feedback to math students is a resource-intensive bottleneck in STEM education. We present Mentir-AI, a tool designed to elevate teacher capacity by generating high-quality mathematical feedback using the Mathforum's "Problem of the Week" archive. By analyzing a corpus of nearly one million interactions, we compare the efficacy of Retrieval-Augmented Generation (RAG) and Fine-Tuning (FT) architectures. This study details the development of an automated grading pipeline, the evolution of a multi-component system prompt, and the implementation of an automated mentor grading system in AI-led evaluation. While Fine-Tuning demonstrates superior instructional judgement, our results identify persistent failure modes in mathematical accuracy and pedagogical judgement. Consequently, we propose a shift from one-shot prompting to an agentic architecture to partition mathematical reasoning from pedagogical drafting.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC