AI systems are increasingly governed by natural language rules, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity. As in legal systems, ambiguity arises both from how these rules are written and how they are applied. But while legal systems use institutional safeguards to manage such ambiguity, such as transparent appellate review policing interpretive constraints, rule-based AI alignment pipelines lack comparable protections. Different interpretations of the same rule can lead to inconsistent model behavior. Drawing on U.S. legal theory, we identify key gaps in current rule-based alignment pipelines by examining how legal systems constrain ambiguity at both the rule creation and rule application steps. We then propose a computational framework that formalizes interpretive ambiguity as constrained entropy minimization and introduces two law-inspired mechanisms: 1) a rule refinement pipeline that minimizes interpretive disagreement by revising ambiguous rules (analogous to agency rulemaking or iterative legislative action), and 2) prompt-based interpretive constraints that reduce disagreement during rule application (analogous to legal canons that guide judicial discretion). We evaluate our framework on a 5,000-scenario subset of the WildChat dataset and show that both interventions significantly improve judgment consistency across a panel of reasonable interpreters. Our approach offers a step toward systematically managing interpretive ambiguity, an essential step for building more robust, law-following AI systems.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex