Summary
In this paper, the authors focus on the generalization performance of linear least squares regression and show the existence of double descent generalization curve in the under-parameterization regime.
In particular, the authors argue, in the linear model in (1) under study (which is slightly different from standard linear models in the literature, but well motivated), that the generalization risk can have a peak in the under-paramererized regime that is due to the alignment between singular space of data and the target, and spectrum of data covariance, instead of the raw dimension ratio or the explosion of the estimator norm.
Some numerical results are provided to support the theoretical analysis.
Strengths
The problem under study is of significance.
The message of this paper looks interesting. But I find it a bit hard to really understand the results and contribution, see my comments below.
Weaknesses
While this paper looks interesting, it is a bit hard for me to really understand the results and contribution.
I think making precise the dimension settings (relation between $n, d, n_{tst})$ will address this issue, see some of my detailed comments below.
Another issue is the contribution: while Theorem 1 is rather general, the discussion thereafter seems all special cases: For example Theorem 2 is a special case, and the results in Sec 4.3 and 4.5 are essentially numerical. The discussions in Sec 5 is interesting but again a very special setting (mixture model of multivariate Gaussian and a fixed direction) without any motivation. It is thus difficult for me to evaluate the significance of this work.
Questions
* line 35: when summarizing the contribution of this work, it would be helpful to forward point to the corresponding theoretical result and/or definition, for the sake of a precise statement of the technical result or the definition (for example, the spiked covariate model).
* it seems that the footnotes 3 and 4 are missing?
* Equation after line 123: I am a bit confused here. What is the purpose here? Is $\hat \beta$ still the min norm solution, then what is $\beta$?
* To make Definition 1 more rigorous, perhaps say here that the convergence of ESD holds in a weak sense as $k \to \infty$, or something like that?
* Theorem 1: perhaps say somewhere in the theorem that this result holds in the asymptotic setting as $n,d,n_{tst}$ going to infinity at the same pace?
* Honestly, I do not understand this result. It seems to me that my previous comment is wrong, and that the result in Theorem 1 does NOT hold in the limit of $n,d,n_{tst} \to \infty$ together, or at least, $n_{tst}$ and $n$ can be much larger than $d$. Some specifications and discussions are needed here.
* Theorem 2 looks interesting. Could the authors comment more on this? For example, note that taking $\mu = 0$ is (more or less) similar to the ridgeless case in the literature. There, according to Theorem 2, we should have a at $c = 1$, as in accordance with "classical" double descent. So, should we understand Theorem 2 as an extension of "classical" double descent to the regularized setting? Or is this due to the model in (1) and (2)?
* line 207 -208: $\hat \Sigma^T \hat \Sigma = \Sigma^T \Sigma + \mu^2$ a typo here?
---
I thank the authors for their detailed reply, which helps me better understand their theoretical results and their contribution I increase my score accordingly.
I, nonetheless, feel that there are a few typos that need be fixed and clarifications needed, in the current version of the paper.
Limitations
I do not see any potential negative social impact of this work.