Summary
The paper examines how information about language and the country of residence influences the replication of a multi-country experiment on political persuasion and mobilization.
To this end, the paper relies on data from a political experiment involving ~7k participants from 15 European countries who read prompts about financial issues in their countries. The prompts attributed blame for the issues to either elites, immigrants, or both, and participants' persuasion and mobilization levels were assessed.
The current research replicates this study, adding participants' demographic information and using an LLM to predict responses. Three experimental settings are tested where the prompts vary by including individuals' country of residence and modifying the language to compare the LLM's performance across conditions.
Reasons to accept
Incorporating and examining cultural information embedded in LLMs is a crucial task, our field increasingly needs accurate approaches for capturing cultural representations and the role of language in shaping them. The paper has this valuable motivation.
This paper makes a valuable point by investigating the correlation between various linguistic factors (e.g., distance from English) and cultural representation, exploring whether and how language influences cultural representation in LLMs.
The research design is noteworthy for its reliance on human data from diverse countries, adding validity to the study of this important question.
Reasons to reject
The paper's approach to assessing the fidelity of LLMs in simulating human responses raises concerns. The chosen method, focusing on replicating regression results, seems inadequately justified. A simpler correlation analysis might have provided a more direct measure of how well the LLMs simulate individual responses.
Additionally, the paper omits mentioning that the data and experiments are limited to European countries and languages. While this doesn't necessarily invalidate the findings, it's a surprising omission that should have been acknowledged.
The generalization of findings to psychological traits is also problematic, as the dataset specifically examines the impact of two contexts on agreement with political arguments about a financial issue. A more appropriate dataset would have included psychological traits collected across diverse cultures, such as those found in the World Values Survey. Furthermore, while the paper discusses culture, it neglects to address the limited cultural variance within European countries.
Questions to authors
The evaluation method assesses 44 coefficients, 30 of which are country-level (C_i^P, C_i^M). However, the LLM is evaluated on replicating these coefficients even though the country of residence is not provided in the prompt. Could you clarify whether the expectation is for LLMs to infer this information from the language and text of the prompt, and if so, what linguistic cues or patterns the LLM might be utilizing for this inference?
The relative deprivation appears to be a crucial independent variable in Bos et al.'s experiment. However, your paper mentions that "the three relative deprivation ratings in our prompt were set to the same average of the actual three relative deprivation ratings by the human participant in the original study" to enhance fidelity. Does this adjustment imply that LLMs struggle to capture the association between relative deprivation and persuasion (and mobilization), and if so, what are the potential reasons for this limitation and how does that impact the paper's main goal for replicating the human responses?