We thank the reviewer for their thorough evaluation of our manuscript and for their constructive comments. We are pleased to address the questions and concerns raised.
1. Most of our results are not self-averaging with respect to a random teacher vector $\boldsymbol{B}$. The generalization error depends on the order parameter $T^{(1)} = \boldsymbol{B}^\mathrm{T} \boldsymbol{\Sigma} \boldsymbol{B}$ (inline equation above Eq.~(5)), which is not self-averaging. This lack of self-averaging arises from the power-law distribution of eigenvalues in $\boldsymbol{\Sigma}$ implying that the order parameter is dominated by a relatively small number of components of the teacher vector. Therefore, we have explicitly averaged over the teacher vectors in the results presented in our manuscript.
2. Regarding the CIFAR-5M plot in Figure 3, we used labels generated by a teacher network. The question about scaling laws with the true labels is intriguing. However, the classification performance of committee machines on CIFAR-5M is not sufficient to enter a scaling regime. We believe that training a more powerful architecture, such as a ResNet, on CIFAR-5M with the true labels would be necessary to observe scaling. We expect that the scaling in this case could exhibit a larger exponent than predicted for our setup, due to the training of multiple layers (see the new plot in Figure 25 of Appendix F in our revised manuscript). In general, scaling exponents depend on the network architectures. For CIFAR-10, Rosenfeld et al. [1] used wide residual networks (WRN) and found a scaling exponent $\beta \approx 0.5$with the model size, whereas Sharma and Kaplan [2] found $\beta \approx 0.23$ for a simple CNN.
To provide additional information on how the plot in Figure 3 was obtained: We used the parametrization from Eq. (80) to produce the results shown in the right panel. We tested our prediction for the generalization error from Eq. (79) on a student network trained on CIFAR-5M images using approximately $10^6$ input examples (see Nakkiran et al., 2021 for details on the dataset). We used only the first channel of the images, resulting in a total input dimension of $N = 1024$ after flattening. To approximate the true covariance matrix $\boldsymbol{\Sigma}$, we numerically estimated the feature-feature covariance matrix based on input examples from the training dataset. During training, we updated only the first $N_l$ entries of $\tilde{\boldsymbol{J}}$, resetting the remaining entries to their initial values after each iteration. Based on the spectra of the feature-feature covariance matrix depicted in Figure 14, we estimated $\beta \approx 0.3$ and used this spectrum to evaluate Eq. (80).
3. In response to the request about plotting Figure 6 on a log-log scale, we have replotted the left panel accordingly. We would like to point out that the inset of this plot was already on a log-log scale in the original version of our manuscript, allowing us to observe the dependence of the scaling exponent on $\beta$. Additionally, in response to another reviewer's request, we have added Figures 22 and 24 to the appendix, displaying the solution of our differential equations on a log-log scale for both error function and ReLU activation functions.
4. Regarding whether the committee machine learns features in the linear network case, for a linear activation function, it cannot learn features. As argued at the beginning of Section 4.1, the teacher output is fully determined by the average teacher vector, i.e., the student does not have access to the outputs of individual teacher units and cannot learn the features encoded by them.
5. Concerning the sensitivity of the solutions in the committee machine to the initialization of the student weights, as described by Eq. (12), the variance with which the student is randomly initialized determines the length of the symmetric plateau. According to Eq. (15), the length of this plateau depends on the teacher's initialization, not on the initialization of the student. Therefore, we expect the dependence on the student's initialization to be weaker than the dependence on the teacher's initialization and to have little effect on the asymptotic power law. In our numerical experiments, we have simultaneously changed the initialization of both the teacher and the student when performing averages.
Our differential equations show that the plateau length depends on the initial conditions, specifically on the $Q^{(l)}$, which are not self-averaging. However, the fluctuations in the plateau length are less pronounced compared to those arising from different initializations of the teacher.
We hope these clarifications and additions adequately address the reviewer's concerns. We have revised the manuscript accordingly and appreciate the opportunity to refine our work based on the valuable feedback.