Traditional evaluations of multimodal large language models (LLMs) have been limited by their focus on single-image reasoning, failing to assess crucial aspects like contextual understanding, reasoning stability, and uncertainty calibration. This study addresses these limitations by introducing a novel benchmark that uniquely integrates multi-image reasoning tasks with rejection-based evaluation and positional bias detection. We further introduce entropy as a novel metric for quantifying reasoning consistency across reordered answer variants and abstention rate to measure a model’s tendency to avoid uncertain answers. We assess 8 state-of-the-art models, such as Grok 3, ChatGPT-4o, Gemini 2.0 Flash Experimental, and DeepSeek’s Janus models, across eight visual reasoning tasks. Our findings reveal ChatGPT-o1 leading in overall accuracy (82.5%) and rejection accuracy (70.0%), closely followed by Gemini 2.0 Flash Experimental (70.8%). On the other hand, Janus models exhibited challenges in bias and uncertainty calibration, reflected in low rejection accuracies (Janus 7B: 25.0%, Janus 1B: 30.0%) and high entropy scores (Janus 7B: 0.839, Janus 1B: 0.787), underscoring their susceptibility to positional bias and unstable reasoning. By employing multi-image contexts, rejection mechanisms, and entropy-based consistency metrics, this benchmark sets a new standard for evaluating multimodal LLMs, enabling a more robust and reliable assessment of next-generation AI systems.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlexReferences (21)
Scroll for more · 9 remaining