Summary
A recent line of work has derived excess error rates for kernel ridge regression in a source and capacity setting under the assumption of Gaussian universality of the kernel features. This work investigates the validity of this assumption in this context.
The main result is that while the rates derived under Gaussian design are correct in the strong regularization regime, the excess error rate can be faster in the weak regularization regime.
Strengths
This is a solid work. The manuscript is clearly written: the context is well explained and the reading is smooth. The results fit in an established literature studying excess error rates for kernel ridge, which recently has seen a revival of interest in the context of deep learning (NTK) and neural scaling laws. Therefore, I believe it is of significant interest to the theoretical community in NeurIPS.
Questions
- In the introduction, the authors say their work addresses three questions. While Q1 and Q2 are precisely addressed by the results, I find that the answer to Q3 falls short. First, the assumption (GF) is surely weaker, but it is still constraining. Second, some of the results in Table 1 are only upper bounds. I understand that most of the excess rate results in the classical kernel literature are also upper bounds, and the authors explicitly discuss that under (GF) it is not possible to derive a matching lower bound, but there is nothing telling us that the picture is not richer in these regimes. Perhaps my problem is with the phrasing of Q3, which differently from Q1 and Q2 is vague.
- I miss a discussion on the intuition behind the (GF) assumption. For instance, in the Gaussian design approximation, one possible intuition is the identification of "orthogonality" (of the features) with "independence". Do you have an intuitive understanding of the two conditions in L440? Why does strong regularization justifies independence?
- Related to the question above, how (GF) differs from the concentration assumption in previous work, e.g. (a1, a2, b1) in [36]. Note that the formulas derived under similar concentration conditions from [36] allow to recover exactly the Gaussian design rates from [16], see the contemporary recent work [DLM]. This suggests (GF) is strictly weaker?
- Since the main result in this work dialogues with previous literature, I suggest commenting and comparing the important notation (source, capacity, regularization decay, etc) in this work with the ones employed in the relevant Gaussian design literature, e.g. [10, 16, 35, 44]. For example, a table like Table 1 in [44] or Table 2 in [16].
- Can you please elaborate on Remark A.5?
[DLM] https://arxiv.org/abs/2405.15699
**Minor points**:
- The authors discuss the "over-parametrized" and "under-parametrized" regime in the text, but never define what they mean. While this can be inferred from the text, it would be good to precisely define it, since this terminology is used in different ways in the ML theory literature.
- For the sake of completeness, it would be better to add the definition of the source in (L106-L108) to the main text in a final version.
- Unpublished pre-prints in the bibliography are missing the arXiv identifiers.
- L110, define $\succeq$ in the notation section.
- L597, maybe $\psi_{k}$ and $\phi_{k}$ are switched?
- L634, "*Consider a kernel $\kappa:\mathcal{X}\times\mathcal{X}\to\mathbb{R}$ be a kernel with [...]*"
- Eq. below L651, missing right bracket.
- Assumption "Domain Regularity (DR)" in Appendix B.2 there are two items (i), (ii) but (iii) is mentioned in the paragraph below (L662-L669) twice.
- L679, precise what "$\lambda$" in $||\bar{\psi}_{i,j}||\lesssim \lambda^{(d-1)/4}$ is.