Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Current frameworks for aggregating free-form text-based opinions with large language models lack the inherent structure needed to provide meaningful fairness guarantees. To address this, we model the task as a multi-objective, token-level Markov Decision Process (MDP), where each objective corresponds to an agent's preference. Each agent's token-level reward is induced by its policy (e.g., a personalized language model). Such policies implicitly define optimal Q-functions, thus enabling stepwise reward computation without an explicit value function. This MDP formulation yields a formal structure that can be analyzed with tools from social choice theory. We first give a stochastic generation policy that is guaranteed to lie in the ex-ante core. It is derived from a distribution over complete statements that maximizes Nash welfare, extending core stability from cooperative game theory and voting to text generation. Second, for a single consensus statement, we target egalitarian welfare and use search within the MDP. Empirically, this search produces statements with improved worst-case agent alignment compared with baselines, including the Habermas Machine. Our code is available https://github.com/cartgr/Generating-Fair-Consensus-Statements-with-Social-Choice-on-Token-Level-MDPs and supplementary material (the appendix) is available https://github.com/cartgr/Generating-Fair-Consensus-Statements-with-Social-Choice-on-Token-Level-MDPs/blob/27abe654c72e465ea58213c80cfb8416573ac2cc/appendix.pdf.

Paper

References (32)

Scroll for more · 20 remaining

Similar papers

© 2026 NYSGPT2525 LLC