While Large Language Models (LLMs) have shown strong capabilities in text-based educational tasks, they struggle with diagram generation due to the absence of visual perception. Vision-language models offer partial solutions but are limited by cost and imprecision. In light of these challenges, we introduce a model-agnostic framework that enables LLMs to reason about and generate diagrams using structured, text-based object representations and callable tools. A custom library defines manipulable diagram elements, while a self-reflective prompting strategy guides the model through iterative planning and revision. Applied to bar modelling in the Singapore math curriculum, our approach produces accurate, interpretable visuals without relying on image understanding—offering a scalable and explainable alternative for visual reasoning in education.
Paper
Full text
Automatic Diagram Generation with LLMs for Visual Reasoning in Education
Semantic Scholar · 2025
Abstract
While Large Language Models (LLMs) have shown strong capabilities in text-based educational tasks, they struggle with diagram generation due to the absence of visual perception. Vision-language models offer partial solutions but are limited by cost and imprecision. In light of these challenges, we introduce a model-agnostic framework that enables LLMs to reason about and generate diagrams using structured, text-based object representations and callable tools. A custom library defines manipulable diagram elements, while a self-reflective prompting strategy guides the model through iterative planning and revision. Applied to bar modelling in the Singapore math curriculum, our approach produces accurate, interpretable visuals without relying on image understanding—offering a scalable and explainable alternative for visual reasoning in education.