VideoMMMU
Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Models scored
26
evaluated
Modality
multimodal
Category
healthcare
+3 more
Published
2025
arxiv.org
Citations
189
Semantic Scholar
Influential
34
citations
References
0
cited works
Venue
arXiv.org
published in
Abstract
Kairui Hu, Penghao Wu, Fanyi Pu, W. Xiao, et al. (+4)
Humans acquire knowledge through three cognitive stages: perceiving information, comprehending knowledge, and adapting knowledge to solve novel problems. Videos serve as an effective medium for this learning process, facilitating a progression through these cognitive stages. However, existing video benchmarks fail to systematically evaluate the knowledge acquisition capabilities in Large Multimodal Models (LMMs). To address this gap, we introduce Video-MMMU, a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos. Video-MMMU features a curated collection of 300 expert-level videos and 900 human-annotated questions across six disciplines, evaluating knowledge acquisition through stage-aligned question-answer pairs: Perception, Comprehension, and Adaptation. A proposed knowledge gain metric, {\Delta}knowledge, quantifies improvement in performance after video viewing. Evaluation of LMMs reveals a steep decline in performance as cognitive demands increase and highlights a significant gap between human and model knowledge acquisition, underscoring the need for methods to enhance LMMs' capability to learn and adapt from videos.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Gemini 3 Pro | 88 | |
| 02 | Gemini 3 Flash | 87 | |
| 03 | Kimi K2.5 | Moonshot AI | 87 |
| 04 | GPT-5.2 | OpenAI | 86 |
| 05 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 85 |
| 06 | Gemini 3.1 Flash-Lite | 85 | |
| 07 | MiniMax M3 | MiniMax | 85 |
| 08 | GPT-5 | OpenAI | 85 |
| 09 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 84 |
| 10 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 84 |
| 11 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 84 |
| 12 | Gemini 2.5 Pro Preview 06-05 | 84 | |
| 13 | o3 | OpenAI | 83 |
| 14 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 82 |
| 15 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 82 |
| 16 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 80 |
| 17 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 80 |
| 18 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 79 |
| 19 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 75 |
| 20 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 75 |
| 21 | Qwen3 VL 8B Thinking | Alibaba Cloud / Qwen Team | 73 |
| 22 | Qwen3 VL 4B Thinking | Alibaba Cloud / Qwen Team | 69 |
| 23 | Qwen3 VL 30B A3B Instruct | Alibaba Cloud / Qwen Team | 69 |
| 24 | Qwen3 VL 8B Instruct | Alibaba Cloud / Qwen Team | 65 |
| 25 | GPT-4o | OpenAI | 61 |
| 26 | Qwen3 VL 4B Instruct | Alibaba Cloud / Qwen Team | 56 |
26 of 26 models · score normalized 0–100 where available