Response to Reviewer WT2A
Dear Reviewer WT2A,
Thank you for your prompt reply.
To address your concern on the GPT-4 generated rationales, we further conducted **human evaluations**. In our previous rebuttal, we performed an automatic evaluation of the entire dataset, and the results indicated the high quality of our data. This human evaluation further proves that the machine evaluators are strongly aligned with human evaluators, confirming the reliability of our automatic evaluations and the high quality of our rationale dataset.
Here are the details: \
**Evaluators:** \
**1) Human evaluators:** We recruited four human evaluators, who are mostly graduate students. They are asked to conduct assessments based on commonsense knowledge and perform Internet searches for validation. On average, it takes them one minute per sample. **2) Machine evaluators:** The latest GPT-4o and GPT-4v models. For each evaluation, we perform three independent runs and calculate the average scores.
**Evaluation Metrics:** \
Factual Consistency: Whether the rationales are consistent with facts \
5 - 100% consistent with fact \
4 - 75% \
3 - 50% \
2 - 25% \
1 - 0% consistent with fact (completely wrong)
Comprehensiveness: Whether the rationales provide all information necessary to predict the category \
5 - cover 100% of discriminative visual features \
4 - cover 75% \
3 - cover 50% \
2 - cover 25% \
1 - cover 0% of discriminative visual features
Visual Disentanglement: Whether the rationales are visually disentanglable or non-overlap \
5 - 100% of rationale visually non-overlap (completely disentangle) \
4 - 75% non-overlap \
3 - 50% non-overlap \
2 - 25% non-overlap \
1 - 0% of rationale visually non-overlap (completely overlap)
**Evaluation Data:** \
We sample **three independent groups** of data from our rationale dataset, each consisting of 50 categories and their corresponding rationales. Specifically, categories were randomly selected from their superclasses: Animals (20), Objects & Artifacts (15), Natural Scenes (5), Plants (5), and Human Activities (5). This ensures not only each superclass is represented but also the robustness of our results [1].
[1] Torralba et al. "Unbiased look at dataset bias." CVPR 2011.
**Quantitative Results:**
| Evaluator | Factual Consistency | Comprehensiveness | Visual Disentanglement |
|:---:|:---:|:---:|:---:|
| GPT-4o | 4.89$\pm$0.05 | 4.55$\pm$0.06 | 4.66$\pm$0.06 |
| GPT-4v | 4.92$\pm$0.03 | 4.67$\pm$0.05 | 4.70$\pm$0.02 |
| **Machine Avg.** | **4.91** | **4.61** | **4.68** |
| | | | |
| Human_A | 4.85$\pm$0.11 | 4.64$\pm$0.19 | 4.42$\pm$0.15 |
| Human_B | 4.97$\pm$0.02 | 4.77$\pm$0.02 | 4.20$\pm$0.11 |
| Human_C | 4.78$\pm$0.04 | 4.60$\pm$0.11 | 4.78$\pm$0.10 |
| Human_D | 4.81$\pm$0.08 | 4.64$\pm$0.05 | 4.77$\pm$0.07 |
| **Human Avg.** | **4.85** | **4.66** | **4.54** |
**Takeaway 1: The quality of our generated rationale dataset is high.** In the table above, we show the results on the average of three sample groups. The average scores demonstrate a high degree of agreement between machine and human evaluators. The dataset consistently achieves scores of 4.61 or higher on the average of machine and human evaluators for each metric, indicating that over 90.3% of the rationales for each category are highly factual, comprehensive, and visually disentanglable.
**Takeaway 2: Our automatic evaluation is reliable and scalable.** As shown in the table above, the average scores of each metric are almost identical between machines and humans. The Pearson correlation coefficient of 0.82 reveals the strong positive correlation between machine and human evaluations. Therefore, using the automatic evaluation method can efficiently evaluate the entire dataset, and align with human evaluations. Note that the results of the entire dataset are reported in the previous rebuttal.
**Qualitative Results:** \
Here we list some categories that rated high in automatic evaluation. As shown, all the rationales are consistent with fact, sufficient for distinguishing the corresponding category, and visually non-overlap (disentangled). \
"ostrich": ["Long, bare neck", "Large, powerful legs", "Small, feathered head", "Two-toed feet", "Black-and-white feathers", and "Long, curved beak"] \
"tabby cat": ["Striped fur pattern", "Dark "M" on forehead", "White paws/chest", "Dark ears/tail tip", "Round face shape", "Whiskers/eyebrows"] \
"airliner": ["Long, slender body", "Multiple engines", "High wingspan", "Narrow fuselage", "Tail fin", "Multiple windows"]
Again, thank you for your constructive comments, we hope this additional human evaluation can address your concerns. Please let us know if our response addresses your questions or if you have any further questions.