Multimodal Coherent Reasoning (MCR): Multistep Multimodal Chain of Thought Based on Large Language Model

Large language models (LLMs) have shown remarkable capabilities in natural language processing, particularly in tasks requiring multi-step reasoning. However, traditional Chain of Thought (CoT) methods, which are primarily utilized for text-based tasks, frequently encounter difficulties in addressing more intricate real-world problems that necessitate multimodal inputs. This paper presents Multimodal Coherent Reasoning (MCR): a new approach to enhance LLM reasoning by generating interpretable rationales from both visual and textual data. The method employs a systematic decomposition of complex questions into sub-questions, utilizing visual models to interpret images and LLMs to generate rationales and sub-answers. Furthermore, the rationales and sub-answers are synthesized into final solutions. Experiments conducted on the ScienceQA benchmark indicate that our method enhances reasoning depth and accuracy. Compared to MM-CoT-Base, our approach demonstrated a 3.37% improvement in accuracy, while DDCoT exhibited a 0.94% increase. Therefore, our method is effective in multimodal reasoning, offering superior generalizability and interpretability

Paper

Full text

PDF

Multimodal Coherent Reasoning (MCR): Multistep Multimodal Chain of Thought Based on Large Language Model

Semantic Scholar · 2024

Abstract

Large language models (LLMs) have shown remarkable capabilities in natural language processing, particularly in tasks requiring multi-step reasoning. However, traditional Chain of Thought (CoT) methods, which are primarily utilized for text-based tasks, frequently encounter difficulties in addressing more intricate real-world problems that necessitate multimodal inputs. This paper presents Multimodal Coherent Reasoning (MCR): a new approach to enhance LLM reasoning by generating interpretable rationales from both visual and textual data. The method employs a systematic decomposition of complex questions into sub-questions, utilizing visual models to interpret images and LLMs to generate rationales and sub-answers. Furthermore, the rationales and sub-answers are synthesized into final solutions. Experiments conducted on the ScienceQA benchmark indicate that our method enhances reasoning depth and accuracy. Compared to MM-CoT-Base, our approach demonstrated a 3.37% improvement in accuracy, while DDCoT exhibited a 0.94% increase. Therefore, our method is effective in multimodal reasoning, offering superior generalizability and interpretability

Similar papers

© 2026 NYSGPT2525 LLC