Summary
This paper analyzed generalization error rates for the Kernel Ridge Regression (KRR) problem. The authors make use of several generic assumptions, including (i) eigenvalue decay, (ii) Embedding index, (iii) source condition, and (iv) Holder continuity of the kernel. It is worth noting that assumptions (i), (ii), and (iv) are generally applicable to popular kernels. The findings of this study align with previous knowledge: in the absence of noise, interpolation achieves optimality, while in the presence of noise, a well-known Bias-Variance tradeoff emerges.
Strengths
1. The paper is well-written, exhibiting a high level of clarity and coherence in its structure and language.
2. The main result presented in the paper is not only insightful but also highly intriguing, contributing valuable insights to the research area.
3. The authors have demonstrated diligent referencing and citation of relevant prior work, showcasing a thorough understanding of the existing literature in the field.
Weaknesses
While I believe the paper is well-written, I feel that there are some areas where the authors assume background knowledge from the reader or make claims without providing adequate supporting discussions. I have identified a few specific points that I believe could benefit from clarification or additional explanation:
1. The contribution section appears to be highly technical. It would be helpful to provide a high-level, less technical overview of the contributions before diving into the specifics. For example, it would be beneficial to provide a brief explanation of the significance of variables such as β and s, as these variables are not defined prior to that section.
2. In line 47, it is mentioned that "our results...MAY suggest the benign...". The term "may" introduces uncertainty, but it is unclear how this conclusion was reached or what factors contribute to this interpretation/uncertainty. It seems that the authors may have intended to convey that since wide neural networks are essentially kernels, and the rates clearly demonstrate that noise impairs generalization, it is unlikely for the noise to have a benign effect. If this is indeed the case, it would be helpful to present this point in a clearer and less convoluted manner.
3. The second row of Figure 1 mentions "overfitting" and "underfitting," but there is no prior discussion or explanation of these terms in the paper. It would be beneficial to define these terms and provide some context.This would enhance the understanding of the results presented in the figure. Overall the second row figures were unclear to me.
4. In line 290, the statement mentions that "These results will help us better understand the generalization mistery of neural networks."
It is unclear how this conclusion was reached or what specific discussions led to it. It would be valuable to elaborate on the implications of the results and provide a more thorough explanation of how they contribute to a better understanding of the generalization of neural networks.(by the way "mistery" is a typo!)
Overall, by addressing these points, the paper can become more accessible to a broader range of readers and ensure a clearer understanding of the key concepts and conclusions.
Questions
In addition to Weaknesses section I have Following questions,
I understand assumption 2 and 3 are common for these kind of analysis. But would you please elaborate,
1. What is the significance of the "embedding index"? What are the implications if α << 1/β or α >> 1/β?
2. What are the limitations of assumption 3? From my understanding, as "s" increases, the function f* becomes smoother since it belongs to H^s. Could you please clarify when this assumption holds or fails?
I am willing to increase my score if the authors address these questions in Weaknesses/Questions sections.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
See weaknesses and questions.