Scaling Laws for Moral Machine Judgment in Large Language Models

Autonomous systems increasingly require moral judgement capabilities, yet whether these capabilities scale predictably with model size remains unexplored. We systematically evaluate 75 large language model (LLM) configurations (0.27–1000B parameters) using the moral machine framework, measuring alignment with human preferences in life–death dilemmas. We observe a consistent power-law relationship with distance from human preferences (D) decreasing as D∝S−0.10±0.01 (R2=0.50, p<0.001) where S is the model size. Mixed-effects models confirm that this relationship persists after controlling for model family and reasoning capabilities. Extended reasoning models show significantly better alignment, with this effect being more pronounced in smaller models (size × reasoning interaction: p=0.024). The relationship holds across diverse architectures, while variance decreases at larger scales, indicating systematic emergence of more reliable moral judgement with computational scale. These findings extend scaling law research to value-based judgements and provide empirical foundations for artificial intelligence (AI) governance.

Paper

Similar papers

© 2026 NYSGPT2525 LLC