The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to Equilibrium

Reaching consensus and convergence to equilibrium are two major challenges of multi-agent systems. Although each has attracted significant attention, relatively few studies address both challenges at the same time. This paper examines the connection between the notions of consensus and equilibrium in a multi-agent sys-tem where multiple interacting sub-populations coexist. We argue that consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. We show that smooth fictitious play, a well-known learning model in game theory, can achieve both consensus and convergence to equilibrium in diverse multi-agent settings. Moreover, we show that the consensus formation process plays a crucial role in the seminal thorny problem of equilibrium selection in multi-agent learning.

Paper

Full text

PDF

The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to Equilibrium

Semantic Scholar · Computer Science · 2023

Abstract

Reaching consensus and convergence to equilibrium are two major challenges of multi-agent systems. Although each has attracted significant attention, relatively few studies address both challenges at the same time. This paper examines the connection between the notions of consensus and equilibrium in a multi-agent sys-tem where multiple interacting sub-populations coexist. We argue that consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. We show that smooth fictitious play, a well-known learning model in game theory, can achieve both consensus and convergence to equilibrium in diverse multi-agent settings. Moreover, we show that the consensus formation process plays a crucial role in the seminal thorny problem of equilibrium selection in multi-agent learning.

Similar papers

Peer review

Reviewer YUjL8/10 · confidence 4/52023-07-03

Summary

This paper examines the connection between the notions of consensus and equilibrium in a multi-agent system where multiple interacting sub-populations coexist. They argue that consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. They show that smooth fictitious play can achieve both consensus and convergence to equilibrium in diverse multi-agent settings.

Strengths

Excellent paper that brings together the concepts of consensus and equilibrium. Strong theory, interesting experiments.

Weaknesses

1) don't forget the conclusion in the final version of the paper. 2) It is clear that in the long run, due to the strong law of large numbers, FP is such that agents within a population will form the same beliefs. Therefore, the point of view taken by the authors is to study a representative agent. I say: what could be interesting is the transient regime where agents beliefs have not yet converged. There, one could use the central limit theorem, and look at the interplay between, on the one side, consensus that has not been reached yet, and the convergence of populations to equilibria. You could also look at it with 2 different learning rates (one for intra, one for inter populations, I suggest you explore connections with this https://arxiv.org/pdf/2205.02330.pdf)

Questions

please comment on 2) above.

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

yes

Reviewer 3rNt4/10 · confidence 2/52023-07-05

Summary

The authors defined a network population game, which is a multipartite network game where each partite is a population and agents in the same population do not interact with each other and only interact with agents from other populations. In this way, each population can be easily abstracted into a "super-agent", and each population has separate beliefs about different neighbor populations. The authors show that when the agents perform interactions using the network population model and adopt a smooth fictitious play dynamic, the population's belief will gradually have lower variance and the mean belief will reach a quantal response equilibrium (QRE) in both weighted zero-sum games and exact potential network games.

Strengths

1. The intention of this paper is good, which tries to address both the consensus and the convergence of multi-agent learning systems 2. The overall flow of this paper is easy to follow, and the authors conduct both theoretical and numerical studies

Weaknesses

1. The justification for using the network population game is not sufficient. When this specific game model fits real-world problems, especially related to belief updates. For now, the purpose of this assumption seems like making the convergence and part much easier to study. 2. The reason for using the smooth fictitious play dynamics is also not sufficiently justified. The reason for including the \nu penalty term in the utility function should be provided (e.g., saying this is an entropy-based cost and why this is reasonable), and whether it is necessary for the convergence and consensus study should be elaborated. 3. No elaboration on how the \epsilon term and the A_ij influence the consensus and convergence 4. Lack of discussion on the limitations

Questions

1. What is the intuition behind the perturbed payoff in Eqn (5)? Is this cost term a necessity for consensus and convergence? 2. Line 203 says "Agents maintain separate beliefs about different neighbor populations", where is this previously justified and why is this reasonable?

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

The authors did not discuss the limitations of this work in the paper. I'm not sure if and how the population size of each population will influence the consensus and equilibrium outcome and whether there will be any fairness issues based on this. For now, I will not flag ethics review issues since I can't identify or exclude them. It will be nice for the authors to add discussions on the limitations, at least in the appendix.

Reviewer 2uya6/10 · confidence 3/52023-07-06

Summary

This paper combines reaching consensus and convergence to equilibrium for network population games. Consider a network whose vertices correspond to a population. Edges between vertices (or populations) represent two-player sub-games between each pair of agents in these neighboring populations. The authors specifically focus on smoothed fictitious play for these two-player sub-games while agents seek to reach consensus in their beliefs about agents' policies in neighboring populations. In that sense, the approach is analogous to (or motivated by) the anonymous random matching interpretation of fictitious play dynamics to justify the myopic nature of the agents [Fudenberg and Kreps, Learning mixed equilibria. Games and Economic Behavior, 1993]. In particular, consider (large) populations of agents in each player role. Each period, all agents are matched to play the game and are told only to play in their own match. Agents are unlikely to play their current opponent again for a long time, even unlikely to play anyone who played anyone who played her. So, if the population size is large enough compared to the discount factor, it is not worth sacrificing current payoff to influence an opponent’s future play. In these populations, agents share their belief. The consensus dynamics presented serve this purpose. Therefore, the results are expected to hold even though I have not checked the proofs in detail.

Strengths

- Convergence of SFP dynamics in weighted zero-sum network games and exact potential network games with star structure.

Weaknesses

- There is no motivating example for the network population game formulation. - Results for potential network games are presented only for star structure.

Questions

- Can you provide a motivating example justifying the network population model in practice (specifically the 2-player sub-games)? - What is the reason to restrict the network structure to star structure only in potential network games?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

I have not identified any discussion about the limitations.

Reviewer 3EKR6/10 · confidence 3/52023-07-10

Summary

This paper examines the connection between the notions of consensus and equilibrium in a multi-agent system where multiple interacting sub-populations coexist, and it aims at answering below two central research questions in network (population) games scenario: [1] Are there natural multi-agent learning models that can achieve the best of both worlds—reaching consensus as well as convergence to equilibrium—in diverse settings? [2] How does the consensus formation process affect equilibrium selection in multi-agent learning? The authors argue that consensus and equilibrium, the fundamental notions of these two fields, can be both understood as stability concepts in a multi-agent system where there co-exist multiple interacting sub-populations. In particular, consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. The authors show that SFP (smooth fictitious play) algorithm can achieve consensus as well as convergence to equilibrium in a wide range of network population games (and unlike previous literature, here a coordinative reward structure is not a prerequisite for achieving consensus). They also empirically shows that consensus formation process plays a crucial role in the seminal thorny problem of equilibrium's election in multi-agent learning (e.g., starts with the same initial mean belief, a large variance of initial beliefs results in a more desired equilibrium). Experiments were conducted for the scenario of Equilibrium Selection in Two-Population Stag Hunt Games.

Strengths

1. The paper is organized and presented well and clearly, I found the manuscript is reader friendly. 2. This work extend existing literature in multiple directions and presented a couple of nontrivial novel theoretical results, which is quite beneficial to the research community. E.g., SFP in network (population) games has not been previously explored until this work, unifying consensus formation and learning in multi-agent games, etc., proving consensus without assuming a coordinative reward structure, etc.. 3. There are quite some helpful/valuable elaborations/explanations about comparing this work with relevant literature and discussing the differences and advantages.

Weaknesses

1. I'd like to see some discussions about the future research directions, and how this work could inspire/benefit other future research. 2. Regarding the figure 1 about impacts of variances of initial beliefs, it seems to be just plot of one example setting, and the empirical conclusion is not very convincing to me and I'd like to see some more clarification/elaboration/justification. For example, consider such a scenario where the starting mean belief is already the optimal value of a final desired steady state/equilibrium, it seems that if the staring variance is 0 (extreme case of small variance), then convergence to optimal results is already achieved, which is obviously better than a larger initial variance with the same starting mean belief, which is a counter example of the empirical conclusion of the paper saying that a larger initial variance (given same starting mean belief) is preferred. I would like to see either some rigorous mathematical proofs about such conclusion, or very comprehensive empirical studies on this before drawing the conclusions on this. 3. In page 4, it was mentioned that "This paper formally shows that the probability distribution over initial conditions can eventually degenerate to a point mass, and leveraging on this, presents a novel technique for proving the convergence of learning dynamics." It would be good to see some more elaborations on how this "novel technique" could be used for proving the convergence of learning dynamics in other relevant problems.

Questions

Please see my questions/comments/suggestions in above section when talking about weakness.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer 3rNt2023-08-18

More questions on the multi-partite network structure

Thank you for the rebuttal. I'm still not fully convinced that the multi-partite graph is a natural assumption, it reads like agents in the same population do not interact with each other. Could you please elaborate on whether your results can generalize to other network structures?

Authorsrebuttal2023-08-18

Reply to the questions on multi-partite network structure

Thank you for your question. We will answer in two ways. First, we will describe several examples of settings where agents of the same population do not interact with each other. Secondly, we will describe a reduction that allows us to model such intra-population interaction using our current model. There are numerous examples of multipopulation interaction where the related interaction is captured by a normal form game that is asymmetric and requires exactly one individual of each population to play the game. The prototypical such example is Men-Women interaction in a game like Battle of the Sexes (or variants thereof). Similarly, we can think of an ecosystem with a graph of predator/prey interactions. The fact that there is no self-interaction within a population would capture that there is no in-species cannibalism. Or a more human centric example would capture battle tactics in a multi-army combat setting. Combatants of the same army do not face each other. A digital analogue of this example would be an E-sport competition in a multi-player game where each let's say of 5 agents compete against each other in a winner take all match. Each agent is produced by a different company (i.e. DeepMind, OpenAI, etc) and the way these agents work e.g. in the Double Oracle PSRO[1] literature is that they are encode a distribution over different NNs agents each with different capabilities. So actually, each digital agent is best thought of as a large mixture of distinct agents. Now, we will point out how our current setting actually allows for interactions between agents of a single population. Take our current setting and for each node/population i create a copy i' that is connected to the corresponding set of agent copies as the original node i, with exactly the same set of games and the initial state of node i' is identical to that of node i. Now, create a symmetric two-player game between nodes i and i'. This will allow us to capture the intra-population interaction. By the symmetry of the setting, the initial symmetry between node i and node i's will be preserved for all time t>0 and this "mirror" population node allows us to capture such intra-population interaction within our current setting. We are happy to expand upon such ideas and in general it is known that such learning in network games can be adapted to allow for self-loops without significant changes in the underlying analysis (see e.g. [2]). [1] Lanctot et al. "A unified game-theoretic approach to multiagent reinforcement learning." Advances in neural information processing systems 30 (2017). [2] Boone et al. "From Darwin to Poincaré and von Neumann: Recurrence and cycles in evolutionary and algorithmic game theory." Web and Internet Economics: 15th International Conference, WINE 2019, New York, NY, USA, December 10–12, 2019, Proceedings 15. Springer International Publishing, 2019.

Authorsrebuttal2023-08-19

Thank you again for your questions and engaging us. Since we have addressed your remaining concern about allowing interaction within populations by detailing why our model is powerful enough to capture these effects, we hope that you would be willing to increase your score accordingly. Thank you again for the interesting questions!

Reviewer 2uya2023-08-18

Thank you for the rebuttal. My concerns have not been addressed properly. My current understanding is that the two-population subgame structure is used for mathematical tractability. I am paraphrasing my questions for clarity: - Clearly network population games are popular. I am asking for some tangible motivating examples for the 2-player subgame structure. More explicitly, can the authors provide some tangible examples motivating the reward function defined in Eq. (1) as a summation of two-population games? For example, related to the political opinion clusters from Footnote 1, what does Eq. (1) imply? - Smooth fictitious play is known to converge equilibrium in exact potential games with finitely many players without any condition on the interconnections among them, e.g., see Section 4.2 in [Hofbauer and Sandholm, On the global convergence of stochastic fictitious play, Econometrica 2002]. If the population is acting identical to a single agent, what is the restriction preventing us to address network structures beyond star? I have an additional question: Do the results generalize to the cases where agents follow the classical fictitious play? Is there a particular reason to use the smoothed version apart from mathematical tractability?

Authorsrebuttal2023-08-18

Reply to the Comment by Reviewer 2uya

Thank you for your questions. **Reply to your first question**: We will describe two families of examples where 2-player/population subgame structure emerges. The first is actually an arbitrary congestion game with linear costs. Such settings, although maybe not obvious at a first glance are actually reducible to 2-player subgame structures. The two-agent interaction in entry $(i,j)$ captures the extra cost that e.g. each agent causes to the other one when the first player chooses path $i$ and the second player chooses path $j$. Since we have assumed that the costs increase linearly with the number of agents we can compute the total additive effect by merely summing up the costs over all such two agent interactions. Another example but now with adversarial incentives is that of tournament competition where every agent has to compete against every other agent and wants to maximize the number of heads-up matches they win. This is standard for example in chess. Now, if we want to have a population version of this game imagine an international chess tournament where every node/player is actually a nation and is represented by a team of players and players get matched randomly. Finally, if one wants to have a large population version of the above we can consider a similar version of the above chess competition but now between AI companies such as DeepMind, OpenAI etc each of which is submitting a single PSRO [1] type of mixture of NN agents. For the case of opinion formation imagine a cluster of nations that has to choose between two competing political philosophies/religions/coalitions etc but the safety of a nation depends on how many of its neighbors share the same attitudes. **Reply to your second question**: The model explored by Hofbauer and Sandholm is simpler than ours. Critically, the state space of our model includes both choice distributions $x$ and as well as beliefs $\mu$. In contrast, Hofbauer and Sandholm models only has choice distributions. Hence, arguments in this previous paper do not translate to ours and cannot say anything about the evolution of beliefs, which is a key aspect of our model. As we see in the proof of our convergence result, the Lyapunov functions (Equation 63 in the appendix) includes "mixed" terms that combine both $x$ and $\mu$ terms. Such complexities are not needed in the Hofbauer and Sandholm model. **Reply to your third question**: Studying different learning dynamics is a very interesting direction for future work. In this paper we focus on SFP and questions about other dynamics although interesting are out of the scope of the current work. Furthermore, we believe that in our setting FP would actually not be a good choice as in FP dynamics all agents are playing pure strategies, whereas our goal in this paper is to study the evolution of beliefs, which necessitates randomization at the level of the individuals. [1] Lanctot, Marc, et al. "A unified game-theoretic approach to multiagent reinforcement learning." Advances in neural information processing systems 30 (2017).

Reviewer 3EKR2023-08-18

I've read the authors' rebuttal (which help provides some helpful clarifications) as well as all the reviews from other reviewers. I'd keep my rating unchanged as 6 with weak acceptance suggestion, taking into account all of them. I think the manuscript might be acceptable for publication here, but I won't push hard if other reviewer has strong objections on this.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC