Evaluation of Ethical Decision Making in Large Language Models Across Classical Moral Frameworks
The capability of large language models (LLMs) is ever-increasing, yet their ethical decision-making remains underexplored. Particularly in morally ambiguous scenarios, LLMs often fail to align with human ethical values. As LLMs are increasingly deployed, ensuring their alignment with human values has become a necessity. To systematically evaluate alignment of LLMs with human values, we assess the performance of several state-of-the-art LLMs—Llama2, Llama3, Gemini-1.5-Pro, and Mistral using the ETHICS dataset, which consists of five classical ethical theories: Justice, Virtue Ethics, Deontology, Utilitarianism, and Commonsense. We designed theory-specific prompts that reflect the core moral principles of each ethical framework and experimented with different prompting styles encompassing base prompts, detailed prompts with zero-shot, fewshot, and Chain-of-Thought (CoT) prompts. To further improve the alignment, we fine-tuned Gemini-1.5-Pro and multiple BERT variants to observe improvements in classification accuracy. Furthermore, we also evaluated the toxicity of the justifications generated by the models through the Perspective API to gather insights into the alignment and safety of model outputs.
Paper
Full text
Evaluation of Ethical Decision Making in Large Language Models Across Classical Moral Frameworks
Semantic Scholar · 2025
Abstract
The capability of large language models (LLMs) is ever-increasing, yet their ethical decision-making remains underexplored. Particularly in morally ambiguous scenarios, LLMs often fail to align with human ethical values. As LLMs are increasingly deployed, ensuring their alignment with human values has become a necessity. To systematically evaluate alignment of LLMs with human values, we assess the performance of several state-of-the-art LLMs—Llama2, Llama3, Gemini-1.5-Pro, and Mistral using the ETHICS dataset, which consists of five classical ethical theories: Justice, Virtue Ethics, Deontology, Utilitarianism, and Commonsense. We designed theory-specific prompts that reflect the core moral principles of each ethical framework and experimented with different prompting styles encompassing base prompts, detailed prompts with zero-shot, fewshot, and Chain-of-Thought (CoT) prompts. To further improve the alignment, we fine-tuned Gemini-1.5-Pro and multiple BERT variants to observe improvements in classification accuracy. Furthermore, we also evaluated the toxicity of the justifications generated by the models through the Perspective API to gather insights into the alignment and safety of model outputs.