Summary
This paper focuses on evaluating the long-context understanding capabilities of multimodal large language models. The authors introduce MILEBENCH, a comprehensive benchmark encompassing multiple dimensions of multimodal long-context understanding, such as Temporal Multi-image Understanding, Semantic Multi-image Understanding, Needle in a Haystack, and Image Retrieval. Additionally, the paper provides a brief evaluation of various multimodal models, ranging from closed-source to open-source, trained on images and videos, presenting a comparative analysis of their performance.
Reasons to accept
1. Long-context capabilities, especially in the domain of multi-image, long-context understanding, are crucial functionalities of multimodal large language models, yet they are absent in many current models. A comprehensive benchmark can facilitate the development of these capabilities.
2. This paper provides an interesting and significant insight: numerous models that excel in common benchmarks show weak performance in multi-images scenario.
3. This paper is well organized, offering concise experiments across a range of popular multimodal large language models.
Reasons to reject
1. This paper collects data from publicly available datasets/benchmarks, yielding only 200 samples for each task. However, some of the pre-existing datasets used herein (e.g., START [1]) encompass a vast collection of video clips and annotations potentially suitable for directly assessing multi-image capabilities. It's not clear whether these datasets alone are inadequate. The authors can provide substantial evidence to justify the necessity of the proposed benchmarks.
2. Recent developments have seen certain models displaying capabilities in multiple images understanding (e.g., MM1 [2], mPLUG-Owl2 [3], MMICL [4]). These models should be included in the evaluation to provide a comprehensive understanding of their capabilities and how training strategy/training data/model designing affect the multimodal long-context understanding capabilities.
3. The experiment results indicate that some models struggle to follow instructions and perform poorly on several samples. It begs the question of whether these models inherently lack the multimodal long-context understanding or if their performance is hindered by unfamiliar patterns that might be improved with training on a few of data samples. In essence, the authors are encouraged to delve deeper into discussing which fundamental abilities, which preventing the models from effectively tackling multimodal long-context problems, are absent.
[1] Wu, B., Yu, S., Chen, Z., Tenenbaum, J. B., & Gan, C. (2021, August). Star: A benchmark for situated reasoning in real-world videos. In Thirty-fifth conference on neural information processing systems datasets and benchmarks track (Round 2).
[2] McKinzie, B., Gan, Z., Fauconnier, J. P., Dodge, S., Zhang, B., Dufter, P., ... & Yang, Y. (2024). Mm1: Methods, analysis & insights from multimodal llm pre-training. arXiv preprint arXiv:2403.09611.
[3] Ye, Q., Xu, H., Ye, J., Yan, M., Liu, H., Qian, Q., ... & Zhou, J. (2023). mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. arXiv preprint arXiv:2311.04257.
[4] Zhao, H., Cai, Z., Si, S., Ma, X., An, K., Chen, L., ... & Chang, B. (2023). Mmicl: Empowering vision-language model with multi-modal in-context learning. arXiv preprint arXiv:2309.07915.