Shaping Mathematics Activities with Generative AI: Prompt Types, Models and Pedagogical Outcomes
This study investigates the relationship between prompt types and the quality of mathematics activities generated by artificial intelligence (AI) tools. Within a multiple-case study design, two advanced AI systems, ChatGPT-5 (OpenAI, September 2025) and Gemini 2.5 Pro (Google DeepMind, September, 2025), were examined using command (C) and request (R) prompts under standardised settings (temperature = 0.7, top-p = 0.9). Four activities were produced and evaluated with the Activity Evaluation and Feedback Tool, which assesses both component-level features (intended outcome, materials, instructions, responsibility, inclusivity, depth, complexity, and mathematical focus) and overall quality. The analysis revealed that three of the four AI-generated activities reached the high-quality range, with total scores of 22, 19, and 23 out of 24 points for Gemini-R, Gemini-C, and ChatGPT-C, respectively, whereas ChatGPT-R scored 15 points, indicating a medium level but close to the high threshold. ChatGPT demonstrated greater effectiveness with command prompts, whereas Gemini produced consistently high-quality outputs, performing better with request prompts. At the component level, intended outcome and materials were consistently strong, while weaknesses were observed in instructions, responsibility, and complexity, depending on the AI–prompt combination. These findings demonstrate that activity quality is shaped not only by prompt design but also by model-specific affordances. Implications are discussed for teacher education, curriculum development, and comparative research on the integration of generative AI in mathematics education.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex