We are glad that you like the paper! The responses to your key questions are below.
**Comparison with other datasets**: We agree that comparing STEM with other datasets will help us better understand the difficulty of different datasets. As the goal of our dataset is to measure the fundamental STEM understanding of neural models, we focus on the evaluation of different models on our dataset. We provide the comparison of statistics of different datasets in Figure 1 (a). This follows the typical benchmark paper design such as VQA [1], MMLU [2], IconQA [3]. However, this is a new angle for the community, and among our future investigations.
**Limitation and impact**: We understand the importance of limitations and potential impact. We have included the discussion in the ethics statement. We plan to expand the discussions to focus on the potential use of models after deep analysis of our dataset. For example, it will follow common limitations of the foundation models such as misuse, bias, and hallucinations [4, 5, 6, 7, 8].
**Figure 9 clarification**: Thank you for pointing this out. Although the general trend is that the model performance increases when models become larger, as the scaling law indicated [9], the model’s performance does not only depend on the size, other factors could also contribute (e.g., the length of the training). It is possible for a model with fewer parameters to outperform another model, which might explain why ViT-B/16 is outperformed by RN50 and RN101.
**Human study**: As described in Sec. 2.4 and Sec. 3.3, our exam scores are calculated based on millions of IXL elementary users [10]. This is one of the main advantages of using our dataset, as the human performances are simulated in a really large environment. The IXL SmartScore [10] serves as the main resource for our human performance comparison. The performance of the graduate students mainly serves as an expert performance, setting up an upper bound for human performance, while the IXL exam scores are from elementary students. We agree that the pool of test takes can be expanded to more people, but it seems a common approach in recent literature such as MATH [11]. We will consider adding more expert test takers in future work.
**More dataset examples**: Due to space limitations, we have included dataset examples for all skills of our dataset in Appendix C and a case study in Section 3.4.
**Directions for improvement**: Thanks for raising this question. We are also considering approaches to improve the performance on the dataset. One possible direction is to include more high-quality textbook data during training since STEM is knowledge-intensive. Another direction could be allowing models to access relevant knowledge resources during inference.
**Ethics statement**: As indicated in our statement, the collected data does not contain sensitive data and the copyright restrictively follows the original data sources. We will add more discussion based on your suggestion.
References:
[1] Antol, Stanislaw, et al. "Vqa: Visual question answering." Proceedings of the IEEE international conference on computer vision. 2015.
[2] Hendrycks, Dan, et al. "Measuring massive multitask language understanding." arXiv preprint arXiv:2009.03300 (2020).
[3] Lu, Pan, et al. "Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning." arXiv preprint arXiv:2110.13214 (2021).
[4] Radford, Alec, et al. "Learning transferable visual models from natural language supervision." International conference on machine learning. PMLR, 2021.
[5] Brown, Tom, et al. "Language models are few-shot learners." Advances in neural information processing systems 33 (2020): 1877-1901.
[6] Ouyang, Long, et al. "Training language models to follow instructions with human feedback." Advances in Neural Information Processing Systems 35 (2022): 27730-27744.
[7] Anil, Rohan, et al. "Palm 2 technical report." arXiv preprint arXiv:2305.10403 (2023).
[8] Touvron, Hugo, et al. "Llama 2: Open foundation and fine-tuned chat models." arXiv preprint arXiv:2307.09288 (2023).
[9] Kaplan, Jared, et al. "Scaling laws for neural language models." arXiv preprint arXiv:2001.08361 (2020).
[10] Learning, I. X. L. "The impact of IXL Math and IXL ELA on student achievement in grades pre-K to 12 (pp. 1–27)." (2019).
[11] Hendrycks, Dan, et al. "Measuring mathematical problem solving with the math dataset." arXiv preprint arXiv:2103.03874 (2021).