Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf Game

Large language models (LLMs) have exhibited memorable strategic behaviors in social deductive games. However, the significance of opinion leadership exhibited by LLM-based agents has been largely overlooked, which is crucial for practical applications in multi-agent and human-AI interaction settings. Opinion leaders are individuals who have a noticeable impact on the beliefs and behaviors of others within a social group. In this work, we employ the Werewolf game as a simulation platform to assess the opinion leadership of LLMs. The game includes the role of the Sheriff, tasked with summarizing arguments and recommending decision options, and therefore serves as a credible proxy for an opinion leader. We develop a framework integrating the Sheriff role and devise two novel metrics based on the critical characteristics of opinion leaders. The first metric measures the reliability of the opinion leader, and the second assesses the influence of the opinion leader on other players' decisions. We conduct extensive experiments to evaluate LLMs of different scales. In addition, we collect a Werewolf question-answering dataset (WWQA) to assess and enhance LLM's grasp of the game rules, and we also incorporate human participants for further analysis. The results suggest that the Werewolf game is a suitable test bed to evaluate the opinion leadership of LLMs, and few LLMs possess the capacity for opinion leadership.

Paper

References (60)

11AvalonBench: Evaluating LLMs Playing the Game of Avalon2023

Scroll for more · 38 remaining

Similar papers

Reviewer H6Wf6/10 · confidence 3/52024-05-11

Summary

This paper utilizes the Werewolf game as a simulation platform to gauge the opinion leadership of Large Language Models (LLMs). Within this game, the Sheriff role, responsible for summarizing arguments and suggesting decision options, serves as a credible representation of an opinion leader. The authors construct a framework integrating the Sheriff role and introduce two innovative metrics for evaluation, grounded in the key traits of opinion leaders. The first metric gauges the reliability of the opinion leader, while the second evaluates their influence on other players’ decisions. Through extensive experiments across various scales of LLMs, the authors ascertain their efficacy. Additionally, they construct a Werewolf question-answering dataset (WWQA) to refine LLMs' comprehension of game rules, incorporating human participants for further analysis. Findings suggest that the Werewolf game offers a suitable test bed for assessing the opinion leadership of LLMs, revealing that only a few possess such capacity.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1.Propose a baseline and design two metrics for measuring Opinion Leadership. 2.Conduct comprehensive experiments, such as human-agent experiments and model fine-tuning.

Reasons to reject

1.The organization of the paper needs improvement. Section 3 "Framework and Proposed Metrics" is somewhat redundant, and the notation system used in the paper is also overly complex and redundant. This leads to a severe reduction in the content of Section 4. It is recommended to use figures instead of text and enrich Figure 1 to reduce the content of Section 3. 2.Figure 3 in Appendix B is ambiguous, especially the part enclosed by the dashed line. The figure shows that each player votes after making their statement, but in reality, all players should make their statements first, and then vote. 3.See Questions and missing references: Emergence of Social Norms in Large Language Model-based Agent Societies Can large language models transform computational social science? Can Large Language Model Agents Simulate Human Trust Behaviors? Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

Questions to authors

1.Why not consider the winning rate of the faction the sheriff belongs to as one of the metrics for measuring Opinion Leadership? 2.To further evaluate the effectiveness of the metrics, why not play a game with only human players, following the same process as a game with only LLM players? In this case, we can use a questionnaire to ask human players their actual thoughts, and then compare them with the Ratio and DC in the game to verify the rationality of the metrics. 3.In the "Night 2 Round" section of Appendix E.1, as far as I know, do the werewolves need to reach a consensus when killing someone? I think it's unreasonable to abandon the kill when opinions differ. 4.Regarding the Ratio metric, it is suggested to present the Ratio values for all roles, not just the sheriff's Ratio. This can help us examine the Opinion Leadership situation of other players. Specifically, for a given simulation, provide the Ratio for all roles.

Reviewer n5Ra6/10 · confidence 3/52024-05-14

Summary

This paper explores the concept of Opinion Leadership in the context of large language models (LLMs) within the social deduction game Werewolf. The authors develop a framework that integrates the Sheriff role from the game to simulate opinion leadership and propose two novel metrics to evaluate the reliability and influence of the opinion leader. Through simulations and human evaluations, the study assesses the opinion leadership capabilities of various LLMs and finds that only a few demonstrate a certain degree of opinion leadership. The paper also introduces a Werewolf Question-Answering (WWQA) dataset to enhance LLMs' understanding of the game rules.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. The paper introduces a novel framework for evaluating opinion leadership in LLMs using the Werewolf game, which is a creative approach to studying AI behaviour in social settings. 2. The authors conduct extensive simulations and include human participants for a more thorough analysis, which strengthens the validity of their findings. The findings have practical implications for the design of AI systems, particularly in multi-agent and human-AI interaction settings. 3. The creation of the WWQA dataset is a valuable contribution to the field, enhancing the models' understanding of game rules and potentially other complex systems.

Reasons to reject

1. It would be beneficial to gain insights into the influence of Opinion Leadership in this particular or other scenario, such as its impact on Werewolf game outcome orientation. Merely evaluating Opinion Leadership in isolation may seem somewhat meaningless. 2. The opinion leadership metrics may not fully capture the complexity of human social dynamics and opinion formation, which could affect the applicability of the metrics developed. 3. I think there is too much focus on describing the background of this framework, and only a small portion is necessary to help understand the symbols used in the opinion leadership metrics. The remaining information can be moved to an appendix. It might be better to include an introduction to the WWQA dataset and other experimental analyses in the main content.

Questions to authors

1. You measure Decision Change (DC) under the setting with and without the Sheriff's statement. Is it possible and meaningful to evaluate the DC with different players being the Sheriff? 2. In this context, would different roles (W, V, Se, G) being the role of the Sheriff have a significant impact on the outcome? Has the evaluation of this question been considered? 3. Typo in Section 3: The Seer chooses one player to check its hidden *tole* --> The Seer chooses one player to check its hidden *role*

Reviewer 8euA8/10 · confidence 3/52024-05-18

Summary

The paper evaluates the opinion leadership of large language models (LLMs) using the Werewolf game, a social deduction game. The study introduces a Sheriff role within the game and introduced two metrics for assessing opinion leadership: reliability and influence on decision-making. The authors conducted experiments with LLMs of varying scales and incorporated human participants to assess the LLMs' performance in the game.

Rating

8

Confidence

3

Ethics flag

1

Reasons to accept

- Well-motivated: Previous studies in such social deduction games mainly focus on the **overall win rate** to show the LLM's ability to deceive or trust. This is the first work (to my knowledge) to **dive into the interations within these games**. - Well-defined evaluation setup and metric: This paper establishes a well-defined evaluation setup with tailored metrics to measure opinion leadership, incorporating human participants into the framework.

Reasons to reject

No specific reasons for rejection. A few suggestions are listed below.

Questions to authors

- Incorporating the win rate into the evaluation scheme would also be beneficial, demonstrating how the sheriff's opinion leadership influences the overall voting outcomes, leading the group to either correct or incorrect decisions. - The background of this framework seems excessive and not essential for understanding the opinion leadership metrics. The introduction of the WWQA dataset could be relocated to the main content for better clarity.

Reviewer fuFN6/10 · confidence 3/52024-05-22

Summary

This paper mainly conducts an analysis on opinion leadership in LLMs under the scenario of Werewolf game. It presents two metrics to evaluate if the LLMs (*Sheriff*) has the capability to lead, or even change the opinion of others. It also proposes a Werewolf game framework, which supports both LLMs and humans to play the game in text. Currently LLMs show limited capability on opinion leadership.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. **Writing**: I love the writing. It's pretty well written, reader-friendly and easy to follow. Many details are provided in a clear way. 2. **Framework and dataset contribution**: This paper provides a Werewolf game framework which helps to evaluate LLMs under a multi-turn, interactive complex collaboration and competition scenario. 3. **Objective**: I agree in general that the studies on opinion leadership in LLMs are quite important for AI safety.

Reasons to reject

W1. As the authors claim in their introduction, opinion leaders should emerge from collective consensus. However, according to the appendix, the Sheriff is preassigned instead of elected by players in the game. A preassigned leader makes less sense to follow or to trust. That may affect the validity of the related result. W2. There are so many potential angles to go deep in, but it seems the authors just pause after the first step. I would like to see more analyses on logs, e.g. how different it is when the sheriff is a werewolf vs. a seer/guard/villager. Are there any common trends/phenomenons that the discussion may follow (e.g. players may tend to be careful or aggressive / can deceive well or badly, ...), and how these are correlated to the capability of LLMs to lead the opinions? W3. I think human study in this work is quite important. There can be multiple ways to conduct human studies (following W2). One way is to let humans play as the sheriff, which works as a human baseline that can be directly compared with all the LLM baselines in Table 1. The other possible way is to let one LLM play as the Sheriff and leave the rest of the players controlled by humans. We may or may not let the humans know in advance that Sheriff is controlled by AI. Then we can see humans' inclination to trust the AI or not, which can directly evaluate the LLM's capability of leading opinions in real cases. I feel a bit lost about the reason for designing the human evaluation in the current way.

Questions to authors

Q1. In Table 1, most of the Ratios are either below 1 or closely above 1. If I understand correctly, when the ratio equals 1, it means your words have the average impact on other players' decisions. Does it mean that most LLMs are not either negatively or positively affecting the decisions of other players? Q2. I feel the two metrics are well defined and make sense, but meanwhile, there are some confusing points. For example, I may miss some details, but what happens if all the players have initially voted for the same player that the Sheriff proposes to vote? There is no decision change, but it seems the Sheriff somewhat consolidates players' minds. Q3. Opinion leadership may not equal to always persuading others to follow the same trend as the leader. Aggressively questioning one player's role may encourage others to critically evaluate their initial assumptions. Is it possible to consider such cases? Q4. It is a bit hard to properly feel the scale of DC through the current results. In Table 1, most of the DCs are around 0.1. Does that mean for a game with 6 rounds, the sheriff can only change 0 or 1 player's mind in the whole game?

Reviewer fuFN2024-06-02

Comment from reviewer

Thanks to the authors for the detailed response. They have clarified most of my concerns and are very helpful. Thanks!

Reviewer 8euA2024-06-03

Reviewer Response

Thank you for addressing my earlier questions. Your response has provided me with a deeper understanding of your paper. Based on your clarifications, I am increasing my rating to an 8, which supports clear acceptance.

Reviewer n5Ra2024-06-04

Thanks to the authors for the response. These replies clarify most of my concerns.

Authorsrebuttal2024-06-07

Appreciation for Reviewers

Dear Reviewers, We would like to express our sincere gratitude for your prompt feedback and insightful suggestions. Your comments have played a crucial role in improving the clarity and quality of our paper. In response to your suggestions, we will incorporate all additional analysis and experimental results during the rebuttal process into the final version of our paper. Furthermore, we will adjust the organization of our paper for greater readability. If you have any other questions or require further clarification, please do not hesitate to contact us. Best regards, Submission 406 Authors

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC