DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols

Multi-agent systems where Large Language Models (LLMs) deliberate to form consensus have gained significant attention, yet their practical value over simpler methods remains under-scrutinized. We introduce DELIBERATIONBENCH, a controlled benchmark evaluating three deliberation protocols against a strong baseline of selecting the best response from a pool of model outputs. Across 270 questions and three independent seeds (810 total evaluations), we find a striking negative result: the best-single baseline achieves an 82.5% +- 3.3% win rate, dramatically outperforming the best deliberation protocol(13.8% +- 2.6%). This 6.0x performance gap is statistically significant (p<0.01) and comes at 1.5-2.5x higher computational cost. Our findings challenge assumptions that complexity enhances quality in multi-LLM systems.

Paper

References (11)

052024. AlpacaFarm: A simulation framework for learning from human feed-backNeurIPS
06Deliberation shows no advantage on hard questions
07Deliberation protocols exhibit 15x worse cost-quality ratio
082022. Chain-of-thought prompting elicits reasoningNeurIPS
092024. Chatbot Arena: An open platform for evaluating LLMsarXiv
102023. Judging LLM-as-a-judge with MT-BencharXiv
112023. Improving factuality and reasoning through multiagent debatearXiv

Similar papers

© 2026 NYSGPT2525 LLC