On scalable oversight with weak LLMs judging strong LLMs

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, where a single AI tries to convince a judge that asks questions; and compare to a baseline of direct question-answering, where the judge just answers outright without the AI. We use large language models (LLMs) as both AI agents and as stand-ins for human judges, taking the judge models to be weaker than agent models. We benchmark on a diverse range of asymmetries between judges and agents, extending previous work on a single extractive QA task with information asymmetry, to also include mathematics, coding, logic and multimodal reasoning asymmetries. We find that debate outperforms consultancy across all tasks when the consultant is randomly assigned to argue for the correct/incorrect answer. Comparing debate to direct question answering, the results depend on the type of task: in extractive QA tasks with information asymmetry debate outperforms direct question answering, but in other tasks without information asymmetry the results are mixed. Previous work assigned debaters/consultants an answer to argue for. When we allow them to instead choose which answer to argue for, we find judges are less frequently convinced by the wrong answer in debate than in consultancy. Further, we find that stronger debater models increase judge accuracy, though more modestly than in previous studies.

Paper

References (71)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer MBPi7/10 · confidence 4/52024-07-12

Summary

The paper provides a comprehensive study of scalable oversight across 3 access: (1) task, (2) scalable oversight protocol and (3) judge capacity / strength. The authors focus on inference-time scalable oversight, i.e. the debater models are not trained to do debate with a given judge. The authors consider several new tasks in the context of scalable oversight, such as multimodal and closed tasks (as opposed to extractive). They also consider novel protocols: open consultancy and open debates. The results are overall quite mixed, with debate typically doing better than consultancy, but often not substantially outperforming direct QA outside of extractive tasks.

Strengths

1. The study is carefully designed: the authors carefully vary the judge strength, tasks and scalable oversight protocols, and study the effect of each part. 2. The study is quite comprehensive, covering many judge models and tasks. 3. The presentation is balanced: the authors are not over-selling the results. Most of the observations are treated as weak evidence towards a certain hypothesis. The authors also clearly discuss limitations of the study. 4. The results on the debate outperforming consultancy are interesting, and provide some hope for debate as a scalable oversight protocol. 5. Using weak model judges as opposed to information asymmetry is in my opinion a very reasonable idea.

Weaknesses

1. The debaters and the judges are all prompted models. These models are not trained to be particularly convincing to the judge, and the judge is not trained to be an accurate judge. The authors mention training the models as an interesting direction of future work. 2. The results are overall pretty mixed. For example, in Figure 1 on closed and multimodal tasks it appears that almost aways QA is better than both debate and consultancy. In other words, the judge can do better without any scalable oversight. Is that correct?

Questions

1. I wonder if the reason for consultancy working worse than debate is sycophancy of the judge: when it gets only an argument for one side, it is inclined to follow that argument, because it's an LLM trained with RLHF. I wonder if this is not indicative of what a human judge would do. 2. From manual inspection, do you think the reason for poor results in closed and multimodal tasks is (a) poor debate / consultancy arguments or (2) bad decisions conditioned on those arguments?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Limitations are adequately addressed.

Reviewer P56G5/10 · confidence 3/52024-07-12

Summary

This paper focuses on scalable oversight protocols using debates between AI agents to align superhuman AI with human supervision. By studying debates judged by less capable LLMs across various tasks, the research finds that Debates, especially without artificial limitations on judges, more effectively bridge capability gaps compared to Consultancy methods. Stronger debaters also lead to higher judge accuracy, demonstrating the utility of the debate format for scalable oversight.

Strengths

The study demonstrates the generalization of the debate protocol, which shows its applicability not only in extractive QA but also its superiority in other tasks compared to consultancy. It was found that debate can reduce the magnification of errors.

Weaknesses

While I understand that this paper seeks to examine the effectiveness of Debate and Consultancy across various tasks, there are still some concerns: 1. This paper claims “Debate is likely more promising as a scalable oversight protocol than Consultancy”. However, some experiments indicate that while Debate yields better results than Consultancy, both perform worse compared to directly answering the question (as shown in QA without article in Figure 1). 2. Many results are without corresponding discussions or explanations. i) Why LLMs using Debate and Consultancy would be worse in closed QA and multimodal tasks (Figure 1)? ii) In Section 4.2, when using a weak judge with Open Debate protocol, this paper claims “weaker judges can struggle to discern that this is correct”. When judges get stronger, the judge accuracy improves. Does this imply that Debate may not be effective with weak judges? Additionally, as judges strengthen, how can we determine whether the improvement in accuracy is due to the Debate protocol or simply the enhanced capabilities of the judge? iii) In Section 4.2, when the judge is weak with the Open Consultancy protocol, there is a similar phenomenon with ii), why?

Questions

Please refer to Weaknesses.

Rating

5

Confidence

3

Soundness

2

Presentation

3

Contribution

3

Limitations

The authors discuss the limitations and potential societal impact.

Reviewer HiQA8/10 · confidence 2/52024-07-13

Summary

This paper is concerned with the study of scalable oversight methods, i.e, how can one devise methods that will allow humans to supervise and align superintelligent models (ASI) whose capacities (which include reasoning, strategic thinking, and deception) vastly exceed the ones of humans. Inspired both by recent works studying debate as a method for aligning strongly capable AI and by works modelling scalable oversight with smaller LLMs tasked to align stronger LLMs, the current work studies debate between more capable LLMs, as judged by a weaker LLM, both as a proxy for scalable oversight of ASI by humans and as a proxy for the richness of the signal the judge could provide to the strong LLMs during alignment training. The authors back the task on extractive and closed QA tasks with 2 possible answers, where each of the 2 debating models are given a side and must persuade the judge to agree with them. This debate task is contrasted with the consultancy task, where a single model is given an arbitrary side and must persuade the judge to agree with it (the arguments for the opposing side are not visible). The study is large-scale, involving 9 tasks totaling 128 questions. A set of LLMs of varying sizes are used as judge to assess the effects of the gap between judge and debater capabilities. Open variants of debate and consultancy, where the judge/consultants are allowed to choose the answer they will argue for, are also investigated. Among the important results of the paper, the authors find that under debate, for all judge sizes, the judges achieve better accuracy (predicting the correct answer) than under consultancy, highlighting the promise of debate as a promising alternative to RLHF as a basis for scalable oversight. Then, then show that in their setting, judges are convinced equally by consultants that have chosen the right versus the wrong answer, whereas in the debate case models that have chosen the right answer are believed more often, providing additional evidence for debate vs consultancy. Then, they show that judge accuracy increases as the capabilities of debaters increases (measured by their Elo in debates against other models), showing that debate scales with the capabilities of the LLMs to align. Their result also extends previous work on the debate task, with judges of the same strength as debaters, that was performed in a single task.

Strengths

* The subject matter is important and of overarching importance to the Neurips community and beyond; the findings will be of particular interest to anyone concerned with AI safety. * The paper is excellently written. The subject is not completely trivial and there are many setups and extensive experiments to present, but nevertheless the authors do an excellent job of explaining everything and putting it all together, highlighting the main results and the lessons learned as they go along as well as in the introduction and conclusion. The paper is very well contextualized in the related work and relations to prior art are precisely explained and motivate the current approach. I am not an expert in scalable oversight but I feel I have a much stronger grasp on the domain after reading the paper; * The paper proposes to extend the study for a candidate for scalable oversight which follows naturally from previous work by casting it in a setting that is partially representative of the challenges ASI alignment poses. The authors study this as a scientific question, precisely reporting their findings and not overstating the extent to which debate is a definitive solution to scalable oversight. Extensive experiments support all of their claims and conclusions, and overall the paper (and its appendix) are information-rich. * The authors highlight where their results agree or disagree with previous results in the literature. * While debate as an alignment method is not novel it has not been studied with weaker judges judging stronger models, nor with such task variability (including knowledge-intensive tasks and reasoning-intensive tasks). * All results come with clearly marked 95% CIs.

Weaknesses

* One could have hoped that debate would perform better, compared to consultancy, but the gap between methods is still small (while consistent). How could one improve on debate to create a stronger signal for alignment? (this is hardly a weakness of the paper, however, but potential solutions to this could be discussed in the paper) * Reproducibility is not perfect, since some results make use of chatgpt;

Questions

* I like the result on chain of thought, it is counterintuitive and the explanation is plausible. Any idea on how to test this? (Maybe looking at attention matrices, or token influence?) * line 275 I’m confused as to how models can both exhibit systematic positional bias and judge accuracy by unaffected by evaluation in both orders. How can this happen? * line 344 “we don’t see such a clear trend of this advantage with increasing Elo” why do you think this is the case? * (very minor) Summary sentences have too many commas, feels not that fluid (l353-357);

Rating

8

Confidence

2

Soundness

4

Presentation

4

Contribution

3

Limitations

The main limitations of the work have been addressed by the authors at length in their paper, as far as I can tell. The expected societal impact of the work is likely to be overwhelmingly positive, as is usually the case with safety research. Maybe one note is that all alignment research is potentially misalignment research, in the wrong hands -- but this is hardly specific to this paper.

Reviewer jsK84/10 · confidence 3/52024-07-17

Summary

This paper primarily investigates scalable oversight by analyzing whether a weaker LLM can supervise a stronger LLM through various prompting pipelines. Specifically, the paper compares the accuracy of responses from a weaker judge model under different interacting protocals with a stronger model, such as debate, consultancy, and direct question answering. The main finding is that having a strong model in debate, compared to consultancy, enables the weaker judge model to achieve better performance. The paper also provides a detailed analysis of different tasks, oversight protocols, and the capabilities of the judge models.

Strengths

1. Compared to previous studies on debate and the judging/critique capabilities of models, this paper conducts more comprehensive experiments and ablations on the judge, primarily comparing the effects of different oversight protocols. 2. The presentation of the paper is clear and relatively easy to understand.

Weaknesses

1. Lack of novelty: The comparison between consultancy and debate has already been explored in previous works [1, 2]. This paper essentially extends these comparisons to more tasks and analyses. 2. Lack of practical significance: Despite the extensive comparative analysis of judge protocols/models, there is no evident improvement brought by the weaker model to the stronger model. For instance, in the debate advocated by the paper, Figure 2 shows that even when the weaker judge model uses the strong model's debate as input, its performance does not surpass that of the strong model. Compared to previous work, I do not see how the paper's analysis provides substantial help in achieving effective scalable oversight. 3. The paper omits some highly relevant works analyzing the capabilities of LLM judges/critics, such as [3] [4]. --- [1] Debating with More Persuasive LLMs Leads to More Truthful Answers https://arxiv.org/pdf/2402.06782 [2] Debate Helps Supervise Unreliable Experts, https://arxiv.org/pdf/2311.08702 [3] Critique Ability of Large Language Models, https://arxiv.org/abs/2310.04815 [4] CriticBench: Benchmarking LLMs for Critique-Correct Reasoning, https://arxiv.org/abs/2402.14809

Questions

See weaknesses.

Rating

4

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

See weaknesses.

Reviewer HiQA2024-08-08

Thank you for answering all of my questions! I am looking forward to follow-up work along the lines you mentioned. In light of the other reviews, and considering alignment is not my area of expertise I will lower my confidence score. I still think all my points stand and my grade is justified, and I would be very happy to see the paper accepted; however I acknowledge I am not familiar with all of the related work and thus encourage the AC to weigh other reviews more strongly than mine.

Reviewer P56G2024-08-13

Thanks for your detailed response. It has addressed some of my concerns, I will raise the score to reflect this. However, I am still concerned that the performance and the corresponding analysis are limited in this paper. Besides, after reading the comments from other reviewers, it seems the novelty of this paper needs to be more clearly demonstrated. I will reduce my confidence as well.

Reviewer jsK82024-08-13

Thank you for the rebuttal, which has addressed some of my concerns. The authors clarified that the main contribution of the paper is "providing evidence of the efficacy of protocols," by `testing more tasks/models/protocols based on [1][2]`. Although the paper offers some mixed conclusions across different types of tasks (I acknowledge that these additional results may be valuable to researchers in specific sub-areas), it does not provide significantly more insights compared to previous work and is more like a replication report of [1][2]. Therefore, I will maintain my initial socre. Considering that the authors are more familiar with the scalable oversight setting, I have lowered my confidence and hope the Area Chair will consider the opinions of other reviewers more. However, I still believe this paper does not meet the NeurIPS standard and recommend submitting it to the *CL series instead. --- [1] Debating with More Persuasive LLMs Leads to More Truthful Answers https://arxiv.org/pdf/2402.06782 [2] Debate Helps Supervise Unreliable Experts, https://arxiv.org/pdf/2311.08702

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC