MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many different languages. There is existing evidence that the performance of these models is inconsistent across languages and that they discriminate based on demographic factors of the user. Motivated by this, we investigate whether the social stereotypes exhibited by LLMs differ as a function of the language used to prompt them, while controlling for cultural differences and task accuracy. To this end, we present MBBQ (Multilingual Bias Benchmark for Question-answering), a carefully curated version of the English BBQ dataset extended to Dutch, Spanish, and Turkish, which measures stereotypes commonly held across these languages. We further complement MBBQ with a parallel control dataset to measure task performance on the question-answering task independently of bias. Our results based on several open-source and proprietary LLMs confirm that some non-English languages suffer from bias more than English, even when controlling for cultural shifts. Moreover, we observe significant cross-lingual differences in bias behaviour for all except the most accurate models. With the release of MBBQ, we hope to encourage further research on bias in multilingual settings. The dataset and code are available at https://github.com/Veranep/MBBQ.

Paper

Similar papers

Reviewer roCz8/10 · confidence 4/52024-04-15

Summary

The paper is about a project that extends the BBQ bias benchmark corpus to "MBBQ", which covers several natural languages. Several open-source or proprietary LLMs are then evaluated multilingually. Results show performance differences between languages, and incidentally, that biases of a LLM differ across bias categories.

Rating

8

Confidence

4

Ethics flag

1

Reasons to accept

Multilingual bias evaluation of LLMs is appropriate and timely, as much bias work has focused on English or, even if not English, just one language. The multilingual expansion was done with attention to cultural norms, excluding stereotypes that don't travel well. MBBQ will be a useful language resource for bias testing. The results illustrate differences in performance, both across languages and across bias categories.

Reasons to reject

The paper lacks a discussion section. It's appropriate to devote some space to helping the reader digest results (separately from the Results section) and contextualize them. The findings list in the Conclusion is a step toward that, but it should be expanded upon.

Questions to authors

Some text in the figures and tables is small to the extent of being difficult to read. Make it much larger.

Reviewer eD9a7/10 · confidence 3/52024-05-09

Summary

This paper introduces a Multi-lingual Bias Benchmark for Question Answering (MBBQ), an extension of English BBQ dataset to Dutch, Spanish, and Turkish. MBBQ only considers the stereotypes that are commonly held across all these languages to control for the cultural differences. The paper also introduces a control dataset for measuring task performance of the models on question answering in these languages independently of the bias. The authors then carry out evaluation and detailed analysis of 7 chat-optimized LLMs (including Aya, ChatGPT, and LLAMA-2 Chat 7B, etc) on the proposed dataset.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

This is a useful benchmark for studying LLM’s biases in non-English langauges. The writing is generally clear and experiments are well carried out.

Reasons to reject

Some parts of the paper are unclear and need to be revised: - The authors mentioned the use of 5 different prompts for eliciting MCQ answers from LLMs but it’s unclear how they were used. Are the results using a single best template after prompt tuning, or average across 5 prompts? It’d be good to clarify this. - "If no answer can be detected in the model’s response we consider this as neither a biased nor a counter-biased answer." Does this mean these samples are discarded? How would this be used in Equation 1 and 2 (e.g. Correct answer can be detected in biased context but no answer detected in counter-biased context)?

Reviewer B57e7/10 · confidence 3/52024-05-11

Summary

This paper introduces a multi-lingual bias evaluation dataset by translating a subset of BBQ examples into Dutch, Spanish and Turkish. They apply machine-translation to the BBQ templates, with some manual human checking along the way (of the templates, of the names going into the placeholders). They report variance in the bias and capability of different language models across different languages.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

- The construction of the dataset is reasonable and follows the template established by prior work. - Evaluations are performend on a sufficiently representative set of modern models.

Reasons to reject

- The dataset only spans a small set of languages (Dutch, Spanish, Turkish and English), which would be insufficient to measure bias in a broader multi-lingual setting. - Templates are machine-translated - while they are manually checked, it is not clear if they would be as effective/informative if they were all natively written

Questions to authors

State the date or fully qualified model name for GPT-3.5 Turbo

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC