Human Misperception of Generative-AI Alignment: A Laboratory Experiment

We conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of risk, time preference, social preference, and strategic interactions, we ask human subjects to make choices for themselves and to predict the choices made by GenAI on behalf of a human user. These problems confront agents with trade-offs (e.g., higher payoff vs. earlier payoff, efficiency vs. equity, riskier but potentially higher rewards vs. safer but lower rewards) and the optimal choices depend on the agent's preferences. We find that people overestimate the degree to which GenAI choices are aligned with human preferences in general (anthropomorphic projection), and with their personal references in particular (self projection). On average, human subjects' predictions about GenAI's choices in every decision environment are much closer to the average human-subject choice than to the average GenAI choice. At the individual level, human subjects' predictions about GenAI's choices in a given environment are highly correlated with their own choices in the same environment. We explore the implications of anthropomorphic projection and self projection in a stylized theoretical model. Our theoretical analysis shows that anthropomorphic projection and self projection can lead to over-delegation to GenAI. More subtly, we also find that objectively improving AI alignment can harm agents who exhibit anthropomorphic projection (because they mistakenly adjust their delegation decisions in a detrimental fashion). Similarly, among agents who exhibit self projection, welfare may be higher for those who have more unusual preferences (since they are less likely to mistakenly delegate). The full paper can be found at: https://arxiv.org/abs/2502.14708

Paper

References (53)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC