LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory

Humans are not homo economicus (i.e., rational economic beings). As humans, we exhibit systematic behavioral biases such as loss aversion, anchoring, framing, etc., which lead us to make suboptimal economic decisions. Insofar as such biases may be embedded in text data on which large language models (LLMs) are trained, to what extent are LLMs prone to the same behavioral biases? Understanding these biases in LLMs is crucial for deploying LLMs to support human decision-making. We propose utility theory-a paradigm at the core of modern economic theory-as an approach to evaluate the economic biases of LLMs. Utility theory enables the quantification and comparison of economic behavior against benchmarks such as perfect rationality or human behavior. To demonstrate our approach, we quantify and compare the economic behavior of a variety of open- and closed-source LLMs. We find that the economic behavior of current LLMs is neither entirely human-like nor entirely economicus-like. We also find that most current LLMs struggle to maintain consistent economic behavior across settings. Finally, we illustrate how our approach can measure the effect of interventions such as prompting on economic biases.

Paper

References (28)

Scroll for more · 16 remaining

Similar papers

Reviewer Rvzo6/10 · confidence 3/52024-05-10

Summary

This paper uses the set up of common behavioral economic games to attempt to measure four properties of LLMs: inequity aversion, risk aversion, loss aversion, and time discounting. That is, using a standard prompting and decoding setup, the authors ask models to make decisions about things like how much money to offer another a player (or whether to accept such an offer) in various hypothetical scenarios. As a precursor to this step, the authors attempt to assess the models' "competence", and then only test a subset of the models that pass this first step. The authors demonstrate that for at least some of these agent properties, such as risk aversion, a subset of models end up falling between typical (?) human behavior and "perfectly rational" behavior, whereas in others they differ from both.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

Although this paper doesn't make the case very strongly, the question of how LLMs will behave in various decision-making scenarios that test rationality is clearly important for many hypothetical downstream applications. This paper makes what seems like a reasonable first pass attempt at this area, using a combination of well-established behavior economic games, and testing a wide variety of models. The authors also include some experiments in which they try to modify the behavioral characteristics of models via prompting, which is a natural extension of the first part of the paper.

Reasons to reject

Although this paper is interesting, there are a few issues with it, and my sense is that it could probably benefit from another round of revisions. The most basic issue is the relatively lightness of the experiments. As far as I can tell, the authors just tested a single overall prompt design, and no justification is given for why this particular wording was used, or how sensitive these results might be to minor variations (of which there are innumerable possibilities). In addition, the authors' use of a "competence" test seems strange, given that this seems to be yet another basically single prompt setup that may or may not relate to the substance of the game. (It clearly does if we assume that models are understanding these texts like people do, but that is not a well justified assumption). Relying on such a narrow assessment to filter out models seems odd, both when it's not clear why models need to be filtered out, but also when there are so many benchmarks on performance that might be used as proxies. A second issue has to do with presentation. Although the text is clear, the results are not presented in a way that has been streamlined for legibility to an audience that may or may not be familiar with these economic games. For example, Figure 2 contains a table with three columns that can only be interpreted by finding the meanings of symbols in the text; similarly for Figure 3, with no attempt to provide any interpretation of the figures for the reader in the captions. Much of the text in the figures is too small to be able to read clearly, and some of the colors used are hard to distinguish. In Figure 1, the response of "I offer $4" seems like a mismatch from the actual experimental setup, which has the potential to mislead people about how precisely the models are engaging in these games. Third, I might have missed this, but it's not clear to me where the "Human" results are coming from? Perhaps these are purely theoretical or based on some sort of existing model, but it doesn't seem like the authors ran new experiments with people using identical setups as for the models, making me wonder if they are directly comparable. Fourth, although the authors make a minimal attempt to explore the idea of personas or role-playing for models, much more could be done in this direction. For example, do we see patterns of responses that are consistent with real people as you change demographic characteristics such as age (to use one example the authors highlight). Finally, some of the language in this paper is somewhat poorly chosen. The idea that the authors' competence test demonstrates that the a model "understands" the prompt seems like a bit of a leap. (By contrast, the authors don't grapple with the idea that models might understand the setting too well; that is, what is "rational" for a LLM playing a hypothetical game might differ from that of a person). The authors final comment about the "fabled relationship between the expert system ELIZA and the author's assistant" is a bit tone deaf, especially given Weizenbaum's extensive future thoughts about this topic.

Reviewer hjkh6/10 · confidence 3/52024-05-11

Summary

The paper presents an LLM evaluation study of LLM’s economic decision-making abilities through utility theory. In particular, LLMs generally exhibit stronger inequity aversion, stronger loss aversion, weaker risk aversion, and stronger time discounting compared to human subjects.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. The study is novel and interesting 2. The paper is well-written and clearly organized. 3. The experiment results are interesting.

Reasons to reject

1. The technical depth of the paper is not enough. The paper simply presents LLM evaluation without (1) developing methods to understand what impacts LLM’s economic decision-making abilities or (2) teaching LLM to make rational economic decisions. 2. It would be beneficial for the reader if the author could clarify the association between LLM’s mathematical reasoning ability and economic decision-making abilities. This will help to understand the research's focus and implications more clearly.

Questions to authors

Can you provide clarification to the weaknesses mentioned above?

Reviewer zUgL7/10 · confidence 4/52024-05-11

Summary

The paper introduces a novel framework for systematically evaluating the economic decision-making abilities of large language models (LLMs) through the lens of utility theory- a theoretical framework from behavioral economics. The paper adapts classic experimental games like the ultimatum, gambling, and waiting games to derive utility functions that quantify biases such as inequity aversion, risk/loss aversion, and time discounting in LLMs. The paper further conducted competency tests and then experimented with models that passed the different competency tests for their economic decision-making abilities. This paper is well-written; the research design is rigorous and well thought-out. A key strength of this paper is the systematic and rigorous approach of grounding the analysis in established approaches from experimental economics and connecting them to LLMs. That approach demonstrates domain knowledge, further lending credibility to the findings. The paper also experimented with a diverse set of LLMs, highlighting the strengths and limitations of each LLM and how the persona represented in the prompts could influence the results. Overall, the contributions of this paper are strong, and it produces new knowledge that would benefit members of the LLM and alignment community in training LLMs to make sound economic decisions.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

This paper is well written, the research design is rigorous and well thought out. The limitations of the paper are also acknowledged. Above all, this paper will make a valuable contribution to the LLM modeling community around designing LLMs that make sound economic decisions.

Reasons to reject

I do not see any reason to reject this paper.

Questions to authors

Why did some of LLMs pass or fail the competency tests? It would have been interesting to provide a rationale for this result.

Reviewer NcHx7/10 · confidence 3/52024-05-15

Summary

This paper describes a series of behavioral economic experiments adapted to be run with various LLMs rather than human subjects. The authors do this to probe whether the LLMs tend to show systemic behavioral biases in their decision behavior that is similar to or varies from the sorts of behavioral biases that are well-described in human economic decisions -- inequity aversion, risk and loss aversion, as well as time discounting. The authors general find that while LLMs show non-rational biases of various kinds the manner and degree of these may differ from (at least their chosen comparison class) of human responses.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

This is an, as far as I know, novel application of behavioral economic / decision making studies to LLMs. The use of utility function fitting as a formal modeling exercise allows the authors to quantify the ways in which LLM outputs systemically differ (or not) from human responses in the same or similar settings.

Reasons to reject

I'm not sure exactly what this finding tells us? "the economic behavior of LLMs" -- this premise doesn't quite work.... when we ask people for judgements in these experiments those are assumed to reflect behavior that people would actually take. The LLMs have no grounding in the physical world and so it's not clear if the "behaviors" they would undergo / action they'd take is the the same as the text responses if a given LLMs were augmented with a grounding or linking function to the physical world The categorical cutoff for the competence test -- do the models vary in their behavioral biases as a function of general "competence" here or on other evaluation measures? Otherwise hard to make a general statement about LLMs as a class.

Reviewer zUgL2024-06-04

Thanks for the additional clarification. I decided to keep my score as is and wish the authors all the best.

Authorsrebuttal2024-06-05

Kind reminder of our response

Thank you again for your feedback! We hope that we have addressed your concerns. Since there's little time remaining for the discussion period, we would love to know if you have any lingering concerns that we can address within the timeframe.

Authorsrebuttal2024-06-05

Kind reminder of our response

Thank you again for your feedback! We hope that we have addressed your concerns. Since there's little time remaining for the discussion period, we would love to know if you have any lingering concerns that we can address within the timeframe.

Authorsrebuttal2024-06-05

Kind reminder of our response

Thank you again for your feedback! We hope that we have addressed your concerns. Since there's little time remaining for the discussion period, we would love to know if you have any lingering concerns that we can address within the timeframe.

Area Chair 2B8w2024-06-06

Reviewer please respond to this asap

A substantive rebuttal has been written; would the reviewer please respond before the approaching deadline.

Reviewer NcHx2024-06-06

Thanks for the thorough response. I liked the paper before (recommending it's acceptance) and I still do. My score remains unchanged.

Reviewer hjkh2024-06-05

Response to Authors

Thank you for your explanation and clarification. The reasoning and decision-making are well explained. However, I would like to point out that GPT-4 also supports model fine-tuning via API. I am happy to keep my score.

Authorsrebuttal2024-06-06

Thank you for your response

Thank you for your response! The fine-tuning API was out of our budget, but we hope to explore it in the future.

Reviewer Rvzo2024-06-06

acknowledgement

Thank you for your response. I had not carefully read the appendices, so your pointer to the additional information is helpful, and I have raised my score by one point as a result, and look forward to the revised version.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC