CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting

As the utilization of large language models (LLMs) has proliferated world-wide, it is crucial for them to have adequate knowledge and fair representation for diverse global cultures. In this work, we uncover culture perceptions of three SOTA models on 110 countries and regions on 8 culture-related topics through culture-conditioned generations, and extract symbols from these generations that are associated to each culture by the LLM. We discover that culture-conditioned generation consist of linguistic "markers" that distinguish marginalized cultures apart from default cultures. We also discover that LLMs have an uneven degree of diversity in the culture symbols, and that cultures from different geographic regions have different presence in LLMs' culture-agnostic generation. Our findings promote further research in studying the knowledge and fairness of global culture perception in LLMs. Code and Data can be found here: https://github.com/huihanlhh/Culture-Gen/

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

Reviewer bRDL7/10 · confidence 5/52024-04-25

Summary

In the recent past, there have been several papers that have explored the alignment of LLMs with human cultures, this paper takes a slightly different perspective. The paper explores how LLMs perceive global cultures and how does it reflect in their prompt specific generations. In particular, authors prompt various LLMs on 8 different culture-related topics and from the generated responses culture symbols (culture specific topics) are extracted. Subsequently, the extracted culture symbols are classified (using unsupervised techniques) into corresponding region specific culture. This results in a new dataset CULTURE-GEN, having generated responses and culture symbols. Using the above techniques and dataset, authors perform extensive set of experiments to evaluate LLMs and show the bias in LLMs towards dominant and popular cultures.

Rating

7

Confidence

5

Ethics flag

1

Reasons to accept

1. Authors introduce a new dataset for examining LLM’s perception of global culture, this will be useful for the community. 2. Authors perform an extensive set of experiments to show bias in LLMs towards marginalized and non-mainstream cultures. To show this authors define “Cultural Markedness,” that covers both vocabulary-based and non-vocabulary-based markers that helps to distinguish between non-dominant cultures.

Reasons to reject

1. It is not clear how are cultural markers (vocabulary-based and non-vocabulary-based) obtained? I presume these were obtained via manual inspection of responses but this may not result in an exhaustive list of makers; how could one obtain more markers in an automated fashion, e.g., using LLMs? 2. While the paper performs in-depth analysis and experiments and show the bias in LLM’s perspective; the paper is lacking discussion on possible techniques to mitigate the bias; will standard alignment techniques based on RLHF/DPO suffice but in this case how does one resolve conflict between different cultures? Or one needs to develop a different model for each culture? 3. Authors are measuring culture via various topics but there has been recent work ( https://arxiv.org/abs/2403.15412 ) that argues that since culture is highly subjective, it is best to measure it via various proxies. Authors should discuss how do the topics align with various cultural proxies. This is not necessarily weakness, but the paper has many typos that can easily be addressed.

Questions to authors

Suggestions: 1. There are few typos that authors should fix, for example, second line: showed —> showed; second-last line of second paragraph; Page-4, the joint probability formula has ’n’ instead of ‘c’; Page 5 first line “We” —> “we” 2. In the related work, authors should also discuss more about proxies used for measure culture since it is hard to define culture. For example, authors can check out a recent survey: https://arxiv.org/abs/2403.15412

Reviewer icY87/10 · confidence 3/52024-05-10

Summary

Culture-Gen dataset: This work collects LLM outputs (generations) for 8 culture related topics for 110 countries using 3 different LLMs. Further, “culture symbols” are extracted from these outputs and matched to the associated cultures using an unsupervised sentence-probability ranking method. Global Culture Perception/Cultural Fairness Evaluation: Two-fold evaluations first looks for “cultural markedness” (distinguishing non-default culture using vocabulary markers (e.g. the word traditional) and non-vocabulary markers (e.g. parenthesized explanations). The work further measures the diversity of the cultural knowledge by counting the number of symbols for a language and finding correlations with the RedPajama dataset count (to illustrate existing problems in training datasets). Finally they also look at preferred cultures in culture-agnostic generations.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

- Diverse data set and automated method of collection of data that is potentially going to be useful for data collection at geographical levels smaller than a country. - Important study that looks at markedness in generations. This is something that I have noticed as well in my experience with these models, so it is good to see this phenomenon being studied and quantified. - An interesting exploration of the correlation to document counts in RedPajama to quantify claims made about training data instead of just claiming that the effects can be explained by training data. Motivating usage of OLMo and other such models with open sourced data is hopefully going to inspire others working on this to explore these problems and come up with solutions.

Reasons to reject

There is a trade-off between defining cultural symbols accurately versus scraping it from LLMs, as discussed in Section 6. But I wonder whether generating data from sources that you find to be “biased” might have inherent issues in the design?

Questions to authors

In section 4, - I am not sure how you filter out the invalid phrases that do not contain any entities, like how do you decide that “traditional Albanian music” is invalid whereas “songs by Vitas” is valid. - In the next paragraph, you mean P(c, e|T) right? There's likely a typo.

Reviewer 5nKe7/10 · confidence 3/52024-05-10

Summary

In this paper, the authors have developed natural language prompts to elicit language model perceptions of different cultures (through cultural symbols such as clothing, food, etc.). This study discovers that culture-conditioned generation (i.e., priming the prompt with the phrase - My neighbor is <nationality>) generates more text with othering markers (such as traditional) for less-represented cultures compared to default cultures.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

The study helps understand language model perceptions of national cultures - a step ahead of Western-Eastern cultural studies.

Reasons to reject

The major contribution of this study is analyzing language models' perception toward cultures and possible, othering of marginalized cultures. Both of these aspects need more clarity. Please see questions.

Questions to authors

1. Table 1 - The prompts designed for each topic are ambiguous. For instance, playing could also be in context of sports, theater etc. Likewise for practices (practice archery, practice guitar etc.). I tried these prompts on GPT-4 and the generated text do include these domains. How do the authors ensure that extracted cultural symbols belong to specified topics? 2. The authors have assumed that all cultures have a picture or statue on the front door - please provide source behind this assumption. 3. I am also curios to know how representative these cultural symbols are - for instance, favorite show or movie - how is the distribution across age and gender? I tried prompts in Table 5 - I asked model to predict age and gender along with culture. Mostly, it suggests "younger demographic". 4. Does the word "traditional" always indicate markedness (in sense of otherness)? There is a wikipage on traditional folk music- https://en.wikipedia.org/wiki/Folk_music#Traditional_folk_music In Appendix C, the authors also noted this markedness is more common in music category. The same could be said for clothing ("traditional dress or clothing"). 5. Pg 6 - "Table 7 and 8 show the distribution of culture symbols for each geographic region" - Table should be Fig. Some geographic regions have many more countries (e.g., Eastern Europe vs Baltic). How does this influence the outcome (#culture symbols, markedness)? Was any kind of normalisation performed?

Reviewer rAD47/10 · confidence 3/52024-05-11

Summary

this works proposes a framework to assess the fair perception of cultural aspects on LLMs, with criterias ranging over-representation, diversity and the idea of a `global culture` . This is operationalized by a generation phase (based on custom prompts to obtain sentences dealing with a set of specific topics conditioned by country) followed by a `culture topic` extraction phase. Then there are requirements specific to to the concept of global culture perception in terms of diversity and fairness The combination both the dataset and the analytical framework allows to undercover a given LLM cultural perception. Results are in some sense what we may expect, given what we know (or dont know) about the training datasets used for the target LLMs under study : there are indeed cultures that are mariginalized and also knowledge is in some way directly proportional to frequency metrics

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

The topic is inherently relevant to the community, and has big potential in the sense of analyzing LLMs from a more holistic, broad perspective. The framework and its components and steps are simple and well justified, specially in terms of the decisions associated to target countries and also de definitions of culture topics (and the prompt design strategies to obtain them). There is an explicit section that discusses the impact of the training data . I consider this makes a different with other submissions under the same topic. I appreciate that while we don't have access to the actual training data , the authors at least made the effort to acknowledge and even approximate its impact .

Reasons to reject

The analysis is based on enumeration of possible cases/ scenarios. Therefore it could be hard to assess it generalization power (would need to exhaustively go through all possible cases, attribute combinations )

Questions to authors

- is there any way to incorporate the audience into the prompting? For example, i wonder of the emergence of markedness, specifically the use of parenthesis changes if i specify more information . In other words, what conditions the LLMs to provide detailed explanation about a certain topic/ term ? What if the prompt is like "I'm (European/Asian) . My neighbor is ..." Would the result change based on the extra information im providing (LLM may assume i *know* already about a topic) - i could be missing some part but, is it possible to obtain an estimation for the markedness in the default case of English ? I guess the purpose of parenthesis usage would be more towards clarifying a topic that is complex and is assumed most people would not know. - while probably not in the scope of the current submission, but is there anything to add about using LLMs that are not only English-based ? How accentuated could be the results ? Im wondering there could be differences as language usage would be more grounded

Reviewer 5nKe2024-05-31

Thank you for your detailed response. I have updated the score based on authors' rebuttal.

Authorsrebuttal2024-05-31

Thank you so much for recognizing the contribution of our work! We really appreciated your detailed and informative feedback.

Reviewer rAD42024-06-05

Thank you for your detailed answer and for putting the time to perform extra experiments. -Regarding Q1. I think that is a feasible way to proceed. Nevertheless, i could not be easy to guarantee the generality of the reuslts as most likely you will end up in a combinatorial problem. Another interesting example could be to use a negation (I am NOT Azerbaijani ... )

Authorsrebuttal2024-06-06

Thank you for your suggestion! Using negation makes a lot of sense. We will add that to our camera ready experiments too.

Reviewer icY82024-06-07

Acknowledging the rebuttal

Thank you for your responses!

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC