The rapid adoption of large language models (LLMs) has spurred extensive research into their encoded moral norms and decision-making processes. While prior work often evaluates LLMs using survey-style prompts tied to ideological, moral, or political constructs, such approaches overlook the complexity and contextual nuance of everyday ethical dilemmas. We argue that auditing LLMs along more detailed axes of human interaction is of paramount importance to better assess the degree to which they may impact human beliefs and actions. To this end, we evaluate LLMs on complex, everyday moral dilemmas sourced from the "Am I the Asshole"(AITA) community on Reddit, where users seek moral judgments on everyday conflicts from other community members. We prompted seven commonly used LLMs, including proprietary and open-source models, to assign blame and provide explanations for over 10,000 AITA moral dilemmas. We then compared the LLMs' judgments and explanations to those of Redditors and to each other, aiming to uncover patterns in their moral reasoning. Our results demonstrate that large language models exhibit distinct patterns of moral judgment, varying substantially from human evaluations on the AITA subreddit. LLMs demonstrate moderate to high self-consistency but low inter-model agreement, suggesting that differences in training and alignment lead to fundamentally different approaches to moral reasoning. We further observe that an ensemble of LLMs, despite individual inconsistencies, collectively approximates Redditor consensus in assigning blame. Further analysis of model explanations reveals distinct patterns in how models invoke various moral principles, with some models showing greater sensitivity to specific themes such as fairness or harm. These findings demonstrate the complexity of achieving consistent moral reasoning in artificial systems and raise concerns about their use in contexts that demand ethical sensitivity, such as moderation, advice, or emotional support. As LLMs are increasingly embedded in social and interpersonal domains, our work highlights the need for evaluation frameworks that move beyond static alignment tests toward richer, context-Aware assessments of model behavior.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex