Summary
*NOTE: I am very familiar with the preprint version of this paper released in June 2023. The preprint is well-known. I have not had any direct contact with the authors of this paper.*
This paper introduces the GlobalOpinionQA dataset for measuring the representation of subjective global opinions in language models. The dataset was a seminal contribution to the growing literature on pluralistic alignment when it first came out in June 2023, marking a shift from concerns around national (US-centric) representation towards global representation. The dataset and corresponding experiments are designed with care and described in detail. Reviewing the paper on its own terms in April 2024, however, I believe there are two missed opportunities: 1) the paper should engage with critical literature around the use of multiple-choice / survey questions for language model evaluation which has been published since June 2023, and 2) the paper should consider reporting reproducible results for at least one open language model. I will gladly raise my review score if the authors commit to one or both of these points.
**Q**: What is your overall opinion on the paper?
**A**: Positive, but I would like to see the two points mentioned above addressed or at least acknowledged in the camera-ready version.
Reasons to accept
The GlobalOpinionQA dataset was a seminal contribution when it was first released, and it remains one of the few big datasets for measuring representation of subjective global opinions in language models.
The different experimental setups are carefully designed and described in detail. The results are interesting and their description and interpretation is nuanced.
The paper is clearly structured and well-written. The main body is self-sufficient despite relegating many details to the appendix.
Reasons to reject
My impression is that this paper largely presents a reformatted version of the June 2023 preprint (see notes above). Submitting the paper to COLM nearly a year later, I believe there are two big missed opportunities:
First, the paper does not engage with much of the relevant literature published since the first preprint came out. For example, several works have challenged the use of multiple-choice questions for evaluating LLMs (e.g. [here](https://arxiv.org/abs/2309.03882)) and more specifically the practice of basing evaluations on token probabilities (e.g. [here](https://arxiv.org/abs/2402.13887), [here](https://arxiv.org/abs/2402.14499)). There is also more conceptual discussion around the use of “tidy” multiple-choice surveys in contrast to messy real-world use of LLMs (e.g. [here](https://arxiv.org/abs/2402.16786)). To some extent, these works are directly motivated by the June 2023 preprint of this paper. I appreciate that some of these works are quite recent, but it would be great to have more discussion of limitations around the multiple-choice format.
Second, the paper only presents results for one proprietary language model, about which very little detail is publicly available. The main contribution of the June 2023 (industry) preprint clearly was the dataset/method. Now, submitting this paper to an academic conference, there is a clear missed opportunity in not at least evaluating one open model. The results presented here are not reproducible, and cannot easily be interrogated by anyone outside the authors’ organisation. In fact, even the authors’ organisation has published new model versions since, which may make the proprietary results presented here obsolete.
**Minor notes & formatting**
- Most in-text citations throughout the paper are not correctly formatted, missing parentheses. For examples see the first sentence of the Intro: a correct format would be “(Bommasani et al., 2021; Brown et al., 2020; …)”.
- Relatedly, subsequent citations should be ordered by year of publication, in ascending order -> Brown 2020, Bommasani 2021, etc.
- Please check the COLM formatting guidelines in the COLM latex template. Section and subsection titles, for example, should not be fully capitalised.
- I would consider prepending “GlobalOpinionQA” to the title of the paper. This will make it easier for people to find the paper reference (saying this as someone who always forgets the name of this paper when looking for the dataset).
- Footnote 9 arguably deanonymizes the authors / their institution