Beyond Seeing: Alternative Representations for Image-Dependent Mathematics

Multimodal large language models (MLLMs) are increasingly proposed for tutoring, feedback, and other educational uses, yet they remain brittle on mathematics problems that require interpreting diagrams, graphs, number lines, and other figures. This dissertation investigates whether many of these failures are fundamentally representational: models may have the mathematical knowledge needed to solve a problem, but fail because the visual information is not encoded in a sufficiently faithful form for downstream reasoning. To address this challenge, I study alternative representations of image-dependent mathematical content, including high-quality alt-text, structured textual descriptions, and executable Python reconstructions of the original figure. I first analyze difficult image-dependent middle-school mathematics problems to identify recurring visual bottlenecks in MLLM performance. I then evaluate whether alternative representations improve not only problem solving, but also educationally meaningful tasks such as hint generation, explanation, scaffolding, accessibility support, and related-problem creation. The dissertation contributes both an empirical account of representational failure in multimodal educational AI and a design framework for making visual mathematical content more usable in learning environments.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC