Summary
This paper introduces a benchmark that assesses the predictive reasoning capabilities of MLLMs across diverse scenarios. The benchmark targets three domains: abstract pattern reasoning, human activity prediction, and physical interaction prediction. The paper evaluates current state of the art LLMs on the benchmark.
Strengths
• The paper addresses an important problem.
• The paper addresses each component task and dataset in detail.
• The paper includes state of the art multi-modal LLMs such as LLaVA and InstructBLIP.
Weaknesses
• Comparison to existing benchmarks for multi-modal LLMs is missing: “Perception Test: A Diagnostic Benchmark for Multimodal Video Models, NeurIPS 2023” already proposes a benchmark suite which includes temporal sequence prediction tasks such as tracking and questions on human actions. “SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension, arXiv 2023” contains questions on action recognition, action prediction and procedure understanding.
• There are already many existing datasets for evaluation of each of the component tasks: abstract pattern reasoning tasks: RPM prediction “Raven: A dataset for relational and analogical visual reasoning, CVPR 2019”, human-centric activity task: ActivityNet-QA “ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering, AAAI 2019”, “STAR: A Benchmark for Situated Reasoning in Real-World Videos, NeurIPS 2021”. It is unclear why the proposed data splits are better than existing benchmarks.
• It is unclear from the paper, the difficulty level of each task. For the human-centric activity task task, the paper chooses 309 and 260 video segments from ActivityNet and Charades respectively. It is unclear how challenging these scenarios are. It would be helpful to include non-LLM based supervised baselines to calibrate the difficult of each task. The paper should include more qualitative examples to highlight the difficulty level of each task.
• Eqs 1-6 seem more like decorative math and are hard to parse. Their realizations in page 6 are much easier to understand and are slight variations of existing evaluation protocols.
• It is unclear how Plausibility, Diversity and Specificity are computed exactly.
• For the Multiple Gold Answer Evaluator, it is unclear how exactly the point-based scoring system is implemented.
• For evaluation of ActivityNet captions standard metrics such as BLEU and Rouge should also be used.
• The benchmark could also integrate an “overall” metric for a global ranking across all tasks.
• The paper could also include GPT-4V as it is the current state-of-the-art multi-modal LLM.
Questions
• The paper should include a more through comparison to prior multi-modal LLM benchmarks.
• The paper should explain in more detail why each component sub-task was chosen.
• Many of the evaluation metrics, e.g., Plausibility, Diversity and Specificity, are not described in detail.
Rating
5: marginally below the acceptance threshold
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.