We thank the reviewers for their time and valuable feedback. During the discussion period, we hope to hear more questions or comments from the reviewers for further discussion to strengthen the paper.
> **the technical difficulty of dealing with the non-i.i.d. noise assumption and the role of the left-spherical symmetry assumption**
We would like to emphasize our contributions that the non-i.i.d. noise assumption itself cannot explain the double descent phenomenon and the left-spherical symmetry (LSS) plays an important role for factoring out $\mathrm{Tr}((X^\top X)^\dagger\Sigma)$ term from the expected variance in the non-i.i.d. noise setting. This trace term explains the double descent phenomenon.
- To further elaborate this, we rewrite the expected variance as $\mathbb{E}\_X[\mathrm{Var}_\Sigma(\hat\beta\mid X)]=c^\top b$, where $a=\lambda((X^\top X)^\dagger \Sigma), b=\lambda(\Omega)$, and $c=\mathbb{E}_X[\Gamma^\top a]$.
- For the i.i.d. noise with $\Omega=\sigma^2 I$ (i.e., $\mathbb{E}[\varepsilon_i\varepsilon_j]=\sigma^21_{i=j}$), the vector $b=\sigma^2 \mathbf{1}$ is parallel to $\mathbf{1}$; and thus we have $c^\top b= (c^\top \mathbf{1})\sigma^2$. Here, we have $c^\top \mathbf{1}=\mathbb{E}\_X[a^\top \Gamma \mathbf{1}]=\mathbb{E}\_X[a^\top \mathbf{1}]=\mathbb{E}\_X[\mathrm{Tr}((X^\top X)^\dagger \Sigma)]$ since $\Gamma \mathbf{1}=\mathbf{1}$ ($\Gamma$ is a doubly stochastic matrix).
- For the non-i.i.d. noise setting (with a general $\Omega$), the vector $b$ is not (necessarily) parallel to $\mathbf{1}$, but with the LSS assumption, $c$ is now parallel to $\mathbf{1}$; and thus we can similarly factorize the expected variance as $c^\top b=\bar c (1^\top b)=\frac1n\mathbb{E}\_X[\mathrm{Tr}((X^\top X)^\dagger \Sigma)] \mathrm{Tr}(\Omega)$ for any positive definite $\Omega$ where $c=\bar c\mathbf{1}$ and $\bar c=\mathbb{E}\_X[\frac1n \sum\_i a\_i]=\frac1n \mathbb{E}\_X[\mathrm{Tr}((X^\top X)^\dagger \Sigma)]$.
- To achieve the factorization $c^\top b=\frac1n\mathbb{E}\_X[\mathrm{Tr}((X^\top X)^\dagger \Sigma)] \mathrm{Tr}(\Omega)$ for any positive definite $\Omega$ (for any $b$ with $b_i>0$), it is **necessary for $c$ to be parallel to $\mathbf{1}$**. And this may not be achieved without the LSS assumption in the non-i.i.d. noise setting.
> **the left-spherical symmetry and the double descent phenomenon**
The left-spherical symmetry makes each eigenvalue of $(X^\top X)^\dagger\Sigma$ have an **equal influence** on the risk.
- The double descent phenomenon can be explained by the trace term $\mathrm{Tr}((X^\top X)^\dagger\Sigma)$ (especially by large eigenvalues) from the expected variance under the left-spherical symmetry (LSS) assumption.
- This is because when $p\approx n$ there are many small eigenvalues of $X^\top X$ near 0 and they highly increase $\mathrm{Tr}((X^\top X)^\dagger\Sigma)$ and the expected variance, but in the overparameterized regime $p\gg n$, the eigenvalues of $X^\top X$ are distant from 0, and thus the variance becomes much smaller.
- Under the LSS assumption, the expected variance is factorized as $\frac1n\mathbb{E}\_X[\mathrm{Tr}((X^\top X)^\dagger \Sigma)] \mathrm{Tr}(\Omega)$ where each eigenvalue of $(X^\top X)^\dagger \Sigma$ is weighted with the same $\mathrm{Tr}(\Omega)$.
- However, this is not the case without the LSS assumption. In other words, without the LSS assumption, the small eigenvalues of $X^\top X$ near 0 may not play the similar role as they do under the LSS assumption.