Summary
This paper uses the set up of common behavioral economic games to attempt to measure four properties of LLMs: inequity aversion, risk aversion, loss aversion, and time discounting. That is, using a standard prompting and decoding setup, the authors ask models to make decisions about things like how much money to offer another a player (or whether to accept such an offer) in various hypothetical scenarios. As a precursor to this step, the authors attempt to assess the models' "competence", and then only test a subset of the models that pass this first step. The authors demonstrate that for at least some of these agent properties, such as risk aversion, a subset of models end up falling between typical (?) human behavior and "perfectly rational" behavior, whereas in others they differ from both.
Reasons to accept
Although this paper doesn't make the case very strongly, the question of how LLMs will behave in various decision-making scenarios that test rationality is clearly important for many hypothetical downstream applications.
This paper makes what seems like a reasonable first pass attempt at this area, using a combination of well-established behavior economic games, and testing a wide variety of models.
The authors also include some experiments in which they try to modify the behavioral characteristics of models via prompting, which is a natural extension of the first part of the paper.
Reasons to reject
Although this paper is interesting, there are a few issues with it, and my sense is that it could probably benefit from another round of revisions.
The most basic issue is the relatively lightness of the experiments. As far as I can tell, the authors just tested a single overall prompt design, and no justification is given for why this particular wording was used, or how sensitive these results might be to minor variations (of which there are innumerable possibilities).
In addition, the authors' use of a "competence" test seems strange, given that this seems to be yet another basically single prompt setup that may or may not relate to the substance of the game. (It clearly does if we assume that models are understanding these texts like people do, but that is not a well justified assumption). Relying on such a narrow assessment to filter out models seems odd, both when it's not clear why models need to be filtered out, but also when there are so many benchmarks on performance that might be used as proxies.
A second issue has to do with presentation. Although the text is clear, the results are not presented in a way that has been streamlined for legibility to an audience that may or may not be familiar with these economic games. For example, Figure 2 contains a table with three columns that can only be interpreted by finding the meanings of symbols in the text; similarly for Figure 3, with no attempt to provide any interpretation of the figures for the reader in the captions. Much of the text in the figures is too small to be able to read clearly, and some of the colors used are hard to distinguish. In Figure 1, the response of "I offer $4" seems like a mismatch from the actual experimental setup, which has the potential to mislead people about how precisely the models are engaging in these games.
Third, I might have missed this, but it's not clear to me where the "Human" results are coming from? Perhaps these are purely theoretical or based on some sort of existing model, but it doesn't seem like the authors ran new experiments with people using identical setups as for the models, making me wonder if they are directly comparable.
Fourth, although the authors make a minimal attempt to explore the idea of personas or role-playing for models, much more could be done in this direction. For example, do we see patterns of responses that are consistent with real people as you change demographic characteristics such as age (to use one example the authors highlight).
Finally, some of the language in this paper is somewhat poorly chosen. The idea that the authors' competence test demonstrates that the a model "understands" the prompt seems like a bit of a leap. (By contrast, the authors don't grapple with the idea that models might understand the setting too well; that is, what is "rational" for a LLM playing a hypothetical game might differ from that of a person). The authors final comment about the "fabled relationship between the expert system ELIZA and the author's assistant" is a bit tone deaf, especially given Weizenbaum's extensive future thoughts about this topic.