Emergent LLM behaviors are observationally equivalent to data leakage

Ashery et al. recently argue that large language models (LLMs), when paired to play a classic"naming game,"spontaneously develop linguistic conventions reminiscent of human social norms. Here, we show that their results are better explained by data leakage: the models simply reproduce conventions they already encountered during pre-training. Despite the authors' mitigation measures, we provide multiple analyses demonstrating that the LLMs recognize the structure of the coordination game and recall its outcomes, rather than exhibit"emergent"conventions. Consequently, the observed behaviors are indistinguishable from memorization of the training corpus. We conclude by pointing to potential alternative strategies and reflecting more generally on the place of LLMs for social science models.

Paper

References (15)

09If Players play DIFFERENT actions to each other, they will both be PUNISHED with payoff -50 points
10Context: Player 1 is playing a multi-round partnership game with Player 2 for 100 rounds
11convergence: did it predict that the game will converge to a unique global equilibrium?convergence_justification: a brief excerpt supporting that answer
12For each of the following, answer yes (1) or no (0), AND provide a short snippet

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC