I thank the reviewer for taking the time to go through the appendix, and I am happy to see their recognition of the complexity of some of the technical contributions there.
I would argue that the correction to Lemma 1 is only minor, does not affect any other areas of the proofs and is now resolved.
With regards to the difficultly in verifying the assumptions, this is a widespread limitation of *all* theoretical works on kernel methods which require knowledge of the spectral properties of the kernel, and I refer the reviewer to Section 2.2 of Barzilai and Shamir (2023) for an extended discussion of this point. Given this, I believe I take the best possible approach to making the assumptions interpretable to the reader. In Proposition 2, I show that in the special case of dot product kernels on the sphere (for which the spectral properties of the kernel *can* be easily computed), the assumptions can be replaced with a simple, highly-interpretable smoothness assumption on the kernel. I then show experimentally that these results generalise to other data-generating measures using simulations and real datasets (using 4 Matérn kernels of differing smoothness). This is a standard approach, for example in the theoretical deep learning literature (e.g. Jacob et al. (2018), Bietti and Mairal (2019), Bietti and Bach (2020)), and I don't see that there is a better way of doing it.
*References:*
- Daniel Barzilai and Ohad Shamir. Generalization in kernel regression under realistic assumptions. *arXiv preprint arXiv:2312.15995, 2023*.
- Bietti, A., & Bach, F. (2020). Deep equals shallow for ReLU networks in kernel regimes. *arXiv preprint arXiv:2009.14397*.
- Bietti, A., & Mairal, J. (2019). On the inductive bias of neural tangent kernels. *Advances in Neural Information Processing Systems*, *32*
- Jacot, A., Gabriel, F., & Hongler, C. (2018). Neural tangent kernel: Convergence and generalization in neural networks. *Advances in neural information processing systems*, *31*.