Thank you for your valuable comments and questions and for the overall positive evaluation of our work! Let us address your question separately below
> The analysis in the paper does not account for the errors induced by discretization. For instance, although it is claimed at the beginning of the paper that the theory covers gradient descent, what is really done is for gradient flow. Also, when the kernel
is discretized using the samples, the following assumptions are all made about this discrete kernel, and it is not obvious how it can be connected to the continuous kernel.
Regarding the first discretization of Gradient Flow (GF) to Gradient Descent (GD), you are right that in the paper, we demonstrate results only for GF. To address this, in the revised version of the paper we double-checked that all claims including gradient-based algorithms are made about GF and not GD. However, are loss functionals derived for Wishart and circle model do not assume any specific shape of the profile, and therefore GD profile $h_t(\lambda)=1 - (1-\alpha \lambda)^t$ can be substituted in the functional to characterize GD generalization. We focused on GF due to its simplicity, and because we do not expect significant differences between GD and GF results in the setting of our paper(but we acknowledge the significant difference in other settings, e.g. for non-linear optimization or when convergence-divergence questions are involved).
Regarding the second discretization of the kernel using samples, I believe that we describe jointly the properties of the continuous kernel via mentioning population spectrum $\lambda_l$ and features $\phi_l(\mathbf{x})$, and discrete empirical kernel matrix $\mathbf{K}$ by mentioning how the training inputs $\mathbf{x}_i$ are generated. For Circle model this is done explicitly. For Wishart model, this is done rather implicitly by the statistics of $\phi_l(\mathbf{x}_i)$ without defining separately features $\phi_l(\mathbf{x})$ and distribution of inputs $\mathbf{x}_i$. Yet, such implicit characterization is sufficient for generalization error in eq. (3) to be well defined, and we can even access it experimentally as we describe in a new appendix section G.1.
> Although this is a theoretical paper, it would be helpful to include some experiments, even if they are on synthetic datasets. This helps explain (and justify) the theory.
That is an excellent comment! During the rebuttal time, we have primarily focused on performing numerical experiments validating our theoretical results. Please see our general response for more details and the revised version of the paper for the new experiment results.
> On page 3, when you define $\mathbf{\Lambda}$, what is the number $P$? Is it the exact rank of $\mathbf{K}$ or its numerical rank, or it does not matter too much?
Throughout the paper, we use both finite and infinite values $P$. For example, all the theoretical calculations use the value $P=\infty$, while in the numerical experiments we take large enough but finite $P$ to stay close to the theoretical setting. While accurate and rigorous treatment of $P=\infty$ requires specific settings for it to be well defined, in our theoretical analysis we follow the "physical level of rigor", prioritizing shorter and more accessible derivations in favor of еру accurate mathematical definitions of all the objects involved.
> In the statement of Proposition 1, I do not think $\rho$ has been defined before and I do not think it is a standard notation outside this community. What is the precise definition of it?
Thank you for mentioning this! Indeed, the version of proposition 1 in the submission was not clear, leaving some ambiguity in the status of measures $\rho$. Conceptually, we believe that these learning measures are new objects introduced in our work (which are essentially equivalent to the loss functional itself). In the revised version of the paper, we improved the wording of proposition 1, including the reference to a general expression of these measures through the expectations over the training dataset $\mathcal{D}_N$.
> Can you make your conjecture on page 6 slightly more formal?
Indeed, that would be a great addition to the work. However, at the moment we feel we are not ready to formulate a rigorous version yet. This would require defining a family of kernel/data models and providing the analysis able to cover all models within the family. It is difficult to do at the moment since our current approach relies on explicit derivation of generalization error, with techniques tailored towards to data models we considered. Intuitively, we expect the condition of Theorem 2 regarding the absence of localization on the largest scale $s=\nu$ to be the principal requirement for the model equivalence.