One-on-one tutoring is highly effective, but most schools and universities cannot provide it at scale, required to support large student populations, primarily due to constraints in staffing, cost, and instructor time. Conventional Intelligent Tutoring Systems (ITS) attempted to fill this gap, yet their reliance on hand-crafted rules and scripted dialogues made them costly and difficult to extend. Advances in large language models (LLMs), including multimodal LLMs that process text and images, offer new opportunities for scalable tutoring. This paper presents a multimodal AI tutor. The system integrates GPT4.1 with a chat interface and interactive whiteboard, allowing students to write formulas, draw graphs, and annotate visuals, while the AI interprets and responds with text and graphics. A multi-agent architecture separates dialogue, mathematical validation, and visual generation to improve accuracy and reliability. Prompt engineering aligned the tutor with principles from learning science, using scaffolding and Socratic questioning to guide students rather than supply direct answers. A user study involving students, staff, and external domain professionals solving a statistics problem demonstrated that the tutor provided context-aware feedback, effective visual interaction, and strong pedagogical alignment. While challenges such as hallucinations and occasional sycophancy remain, results demonstrate that multimodal LLMs can approximate the benefits of one-on-one tutoring and should be explored as a complement to classroom teaching.
Paper
Full text
An AI Tutor System with a Virtual Whiteboard for Statistics Learning
Semantic Scholar · 2026
Abstract
One-on-one tutoring is highly effective, but most schools and universities cannot provide it at scale, required to support large student populations, primarily due to constraints in staffing, cost, and instructor time. Conventional Intelligent Tutoring Systems (ITS) attempted to fill this gap, yet their reliance on hand-crafted rules and scripted dialogues made them costly and difficult to extend. Advances in large language models (LLMs), including multimodal LLMs that process text and images, offer new opportunities for scalable tutoring. This paper presents a multimodal AI tutor. The system integrates GPT4.1 with a chat interface and interactive whiteboard, allowing students to write formulas, draw graphs, and annotate visuals, while the AI interprets and responds with text and graphics. A multi-agent architecture separates dialogue, mathematical validation, and visual generation to improve accuracy and reliability. Prompt engineering aligned the tutor with principles from learning science, using scaffolding and Socratic questioning to guide students rather than supply direct answers. A user study involving students, staff, and external domain professionals solving a statistics problem demonstrated that the tutor provided context-aware feedback, effective visual interaction, and strong pedagogical alignment. While challenges such as hallucinations and occasional sycophancy remain, results demonstrate that multimodal LLMs can approximate the benefits of one-on-one tutoring and should be explored as a complement to classroom teaching.