Details of Manual Evaluation and Human Performance
We thank the reviewers for their insightful feedback. We are encouraged that they agreed with us on the timely manner and importance of this topic (R1, R3) and have found our analysis to be rich (R1) and insightful (R2, R3), revealing significant shortcomings in the current open-source MLLMs (R2, R4) and enabling future research (R1, R4). In the following, we address their questions and concerns regarding human performance and manual evaluation:
**Human Performance:**
To address the questions regarding the comparison to human performance, we ran a study with 25 college-level participants from diverse educational and demographic backgrounds. We provided each participant with ten randomly selected samples from the IQ50 dataset (20% of the dataset) and instructed them to solve the puzzle and give a short reason for their answer. The average performance of this group was **95.9%**, with a standard deviation of 6.92%. We will add these results to the paper to further showcase the performance disparity.
**Manual Evaluation:**
Here, we provide more details to address the questions regarding the manual evaluation procedures:
a) Regarding the demographic and expertise of the evaluators and annotators, all were mid-level to senior Computer Science PhD students specializing in NLP with extensive experience working with and evaluating LLMs.
b) Regarding the rubric and the inter-annotator agreement, our main challenge was to overcome the nuances that appear in reasonings generated by LLMs. As such, the evaluators first met to 1) determine the correct reasoning paths for the samples in the dataset and 2) review a series of sample responses generated by the models (e.g., GPT-4v, Gemini, etc.) to determine the evaluation strategy. Based on the observations in this initial meeting, the evaluators decided to allow for extra/wrong details in the generated responses as long as they did not affect or interfere with the alignment of the generated response to correct reasoning paths.
For example, in some instances, the models perceived shadows (or incorrect colors) in the shapes that were not present in the puzzle; however, as long as they correctly detected the row-wise and column-wise change of patterns (e.g., square turning to circle) and grounded their reasoning on them, the evaluators marked the response as correct.
Moreover, each sample was annotated and assessed by one person, and any uncertain case was flagged and shared among all three evaluators for discussion. After discussions, the final label was determined by a majority vote (i.e., at least 2 out of 3).
c) Regarding the scalability of evaluations, while we acknowledge the difficulty of scaling such manual experiments, one of the main challenges of correctly assessing generated responses is that automatic metrics like ROUGE, BERTScore, etc., fall short of adequately evaluating the semantic nuances. Hence, we decided to go down the path of highly labor-intensive human expert evaluation to provide precise, concrete, and grounded insights.
While we acknowledge that manual evaluation for LLM-generated responses is nuanced and challenging, we tried to avoid any pitfall that could invalidate our results (e.g., using only expert PhD students for assessments and annotations). Furthermore, we will release the models' outputs and the evaluators' annotations to mitigate further concerns about the quality of the evaluations.