Multimodal Large Language Models (MLLMs) have significantly advanced vision-language integration, enabling applications such as visual question answering, document understanding, and multimodal reasoning. Despite this progress, existing models still suffer from limitations in reasoning accuracy, interpretability, and cross-modal alignment. To mitigate these issues, the Visual Chain-of-Thought (Visual CoT) paradigm has emerged, extending step-by-step reasoning to multimodal settings by explicitly modeling intermediate reasoning processes grounded in both visual and textual evidence. In this survey, we provide a comprehensive review of Visual CoT. We trace its origins and evolution, and present a unified taxonomy that categorizes existing approaches into text-based and image-based Visual CoT according to the carrier of reasoning. We further refine this taxonomy by grouping text-based Visual CoT by reasoning generation and control, and image-based Visual CoT by visual state manipulation or generation. Finally, we summarize representative benchmarks and evaluation protocols, discuss open challenges such as computational efficiency and modality alignment, and highlight promising directions for future research. This survey aims to establish a systematic foundation for studying reasoning in MLLMs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex