Response to Reviewer P6cQ
We thank the reviewer for the positive evaluation of our paper and the helpful comments on how to improve the paper. The feedback and questions are answered in the order they were brought up in the review.
---
*The model is quite abstract at some places. For the theoretical results, they are mostly about the analysis of the game and I am not sure how relevant they are for this conference (although they are certainly interesting for a certain community). It might have been more interesting to focus more on the learning algorithm.*
A: The analysis of the game provides the key insights into complex agent systems that are necessary to eventually provide the equilibrium learning algorithm. Only through a thorough understanding of the core periphery structure and its implications it is possible to state a principled equilibrium learning approach. Therefore, we do believe that the understanding of these complex agent networks and the resulting learning algorithm are relevant for this conference. Nevertheless, we agree that there are various open challenges and hope that our GXMFG learning approach provides a useful framework for future research.
---
*Q1: I am wondering if some assumptions are missing. For example below Lemma 1, should $f$ be at least measurable (and perhaps more?) with respect to $\alpha$ for the integral to make sense?*
A: We thank the reviewer for pointing out the potential issue with the assumptions on $f$. In the theoretical results such as Theorem 1 we state the assumptions on $f$, i.e. measurable and bounded. However, we did not mention these conditions on $f$ when defining the integral you referred to. To avoid the misleading impression that there are no assumptions on $f$, we also added them to the respective definition in the updated draft.
---
*Q2: Assumption 2 as used for instance in Lemma 1 does not seem to make much sense (unless I missed something): What is $\boldsymbol \pi$? We do not know in advance the equilibrium policy and even if we did, we would still need to define the set of admissible deviations for the Nash equilibrium. Could you please clarify?*
A: We completely agree with the reviewer: the policy $\boldsymbol \pi$ should not be part of Assumption 2 and (of course) we do not assume the equilibrium policy to be known in advance; the set of admissible deviation policies from the Nash equilibrium is not restricted. Instead, we have added the Lipschitz condition (up to a finite number of discontinuities) on $\boldsymbol \pi$ to the respective theoretical results, such as Theorems 1-4. Thank you for spotting the mistake in Assumption 2. We have corrected it in the updated paper version.
---
*Q3: Algorithm 1, line 14: Could you please explain or recall what is $Q^{k, \mu^{\tau_{\max}}}$?*
A: In Algorithm 1, $Q^{k, \mu^{\tau_{\max}}}$ is defined similar to $Q_{i,t}^{\pi, \mu}$, except that we substitute the reward function $r$ by $r'_k$ and use the transition kernel $P'_k$ instead of $P$. We have added the definition in the updated paper.
---
*Some typos: Should the state space be either $\mathcal{X}$ or $X$ (see section 3 for instance)?*
A: The state space should be $\mathcal{X}$. We used $X$ in the beginning of Section 3 to denote an arbitrary finite set. Thanks for pointing out the ambiguous notation; we have corrected it in the updated paper version.
---
*Does $\mathbb{G}^\infty_{\alpha, t}$ depend on $\mu$ or not (see bottom of page 4)? Etc.*
A: In our framework, the neighborhood distribution $\mathbb{G}^\infty_{\alpha, t}$ always depends on the mean field $\boldsymbol \mu$. For notational convenience, we sometimes drop the dependence on $\boldsymbol \mu$ in the notations. The lack of any comment on dropping the dependence in the notation led to understandable confusion. We have added an explanation to the updated draft, thanks.