Generalized equivalences between subsampling and ridge regularization

We establish precise structural and risk equivalences between subsampling and ridge regularization for ensemble ridge estimators. Specifically, we prove that linear and quadratic functionals of subsample ridge estimators, when fitted with different ridge regularization levels $\lambda$ and subsample aspect ratios $\psi$, are asymptotically equivalent along specific paths in the $(\lambda,\psi)$-plane (where $\psi$ is the ratio of the feature dimension to the subsample size). Our results only require bounded moment assumptions on feature and response distributions and allow for arbitrary joint distributions. Furthermore, we provide a data-dependent method to determine the equivalent paths of $(\lambda,\psi)$. An indirect implication of our equivalences is that optimally tuned ridge regression exhibits a monotonic prediction risk in the data aspect ratio. This resolves a recent open problem raised by Nakkiran et al. for general data distributions under proportional asymptotics, assuming a mild regularity condition that maintains regression hardness through linearized signal-to-noise ratios.

Paper

Similar papers

Peer review

Reviewer 96Gw7/10 · confidence 3/52023-07-04

Summary

The authors investigate the problem of ridge regression, and prove equivalence results between ridge regularisation and ensambling of weak learners trained on subsamples of the original dataset. The equivalences hold under very mild assumptions, and notably there is no requirement on the data model. Two kind of equivalence are proven: i) equivalence at the level of a quite generic class of risks ii) equivalence at the level of the ridge estimator itself. The equivalence basically say that one can trade a bit of subsampling for a bit of ridge regression without altering the performance of the estimator. The equivalences hold on paths in the plane defined by the ridge regulariser and the subsampling ratio. The authors provide both a "theoretical" characterisation of such paths, which requires knowledge of the population covariance of the features, and a "data-driven" characterisation, which requires only access to the sample covariance of the features. Finally, the authors discuss possible extensions of their results to real-data scenarios and random features regression.

Strengths

The works seems sound, relevant (answering open questions in the literature), well-motivated and well-presented. Code is available for reproducibility. I would like to highlight that the authors need basically no structural assumptions on the data model to prove the equivalence.

Weaknesses

I did not identify any substantial weakness in the paper.

Questions

I do not have any question of the authors. I only suggest the authors to improve Figure 3: the caption could additionally describe the difference in x-axis between left and right panel.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors do not discuss explicitly limitations. I find that this is not strongly necessary, as the assumptions under which the results hold are clearly presented. The only improvement I could suggest is for the authors to discuss whether they expect the equivalences they prove to break in some specified setting.

Reviewer MtFr10/10 · confidence 5/52023-07-05

Summary

This paper shows an asymptotic equivalence between an ensembled+subsampled (E+S) version of ridge regression and the standard version, in the proportional asymptotic regime. The equivalence result shows that there exists a linear path in the space (ridge-parameter, aspect-ratio) along which all estimators yield essentially the same solution. This equivalence also resolves an important open question regarding the behavior of optimally tuned ridge regularization. Specifically, the generalization error monotonically decreases with overparameterization (assuming the same level of SNR).

Strengths

This paper was delightful to read. The problem considered is of very broad interest. The contribution is fundamental and is potentially path-breaking. At the very least, it yields a very satisfactory understanding of the effect of ridge regularization under both important settings. - optimal tuning - interpolation (lambda = 0)

Weaknesses

The paper can do with a round of proof-reading. There are some minor issues. Perhaps the title should read "ensembled subsampling" instead of just "subsampling". A sketch of the proof is missing. It would be good to highlight in a couple of paragraphs the core ideas behind the proof of the main result.

Questions

In Section 2, does M have to grow at a certain rate wrt n ? It appears that the Conjecture regarding Kernel Ridge Regression may already be within the reach of this paper, following the equivalence result from Sahraee-Ardakan et al. https://arxiv.org/pdf/2201.08082.pdf

Rating

10: Award quality: Technically flawless paper with groundbreaking impact, with exceptionally strong evaluation, reproducibility, and resources, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

No negative impact envisioned.

Reviewer nskG7/10 · confidence 2/52023-07-06

Summary

The authors study the relationship between ridgeless ensembles constructed from subsampled data and a ridge estimator in a setting with mild assumptions on the joint distribution $(Y,X)$. They establish equivalences for a generalized class of risk functionals, which include quantities related to coefficient estimation and both in and out of sample errors. These equivalences are proven in the case where the ensemble includes estimators trained on all possible subsamples of a given size; the authors extend these results to equivalences between two finite ensembles. The authors use these equivalence results to settle a conjecture in a previous paper about risk monotonicity of ridge regression as a function of $p/n$.

Strengths

- The authors substantially relax the distributional assumptions considered by previous work in the literature. This is conceptually important since it was not known how critical the linearity assumption is for these types of equivalences to hold. - The theoretical analysis is both novel and technical, invoking various concepts from random matrix theory. - Overall, the paper is well-written and well-organized.

Weaknesses

While equivalences between ridge regression and subsampling are conceptually interesting, at this stage it appears that consequences for data analysis are a bit limited. However, the authors generously discuss several potential extensions for which some of the tools developed in the paper may be helpful.

Questions

- In Theorem 3, is the ensemble size $M$ fixed? - While the paper is well-written overall, some additional exposition/clarification in certain parts may be helpful to readers. For example, the authors could elaborate more on what they mean by first-order and second-order. In addition, the coefficient confidence interval case can be explained more. Also, before theorem 3, it is stated that "We can go a step further and ask if there exist any equivalences for the finite ensemble and if there are any equivalences at the estimator coordinate level between the estimators." Do you mean that the conditions in this theorem have been previously shown to imply equivalences at the coordinate level and these conditions also imply equivalences with finite ensembles or do you mean something else?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

The authors are quite forthcoming about various limitations of their work and possible future extensions.

Reviewer wEHo6/10 · confidence 3/52023-07-15

Summary

This submission establishes equivalences between ridge regression (i.e. $\ell_2$-penalized linear regression) and ensembles of linear models trained on sub-sampled datasets. In particular, the authors prove that for a fixed feature/sub-sample-size $d/k$, ratio, there exists a ridge-regression model with risk asymptotically equivalent to the full ensemble, i.e. the average of all $k$ models trained on sub-samples of size $k$. Moreover, this equivalence holds for all convex combinations of the ridge-regression and full sub-sampled model. Note that the asymptotics assume $d$, $k$, $n$ approach infinity such that the ratios $d/k$ and $d/n$ are held constant. This equivalence result is then extended to other metrics, such as training error and the weight estimation error, and to "structural" results which show asymptotic equivalence of the weights of the models. The authors leverage these equivalences to show that the prediction risk of the best ridge regression model is monotone increasing in $d/n$.

Strengths

The main strength of this work is the theoretical contributions. In particular: - The authors extend exist results on equivalent risk of ridge regression models and sub-sampled ensembles to new metrics and to the model weights themselves ("structural results"). - The theorems are proved under relaxed conditions compared to previous work. In particular, general data distributions are allowed provided a fairly weak assumption on the moments holds. - The authors answer an open problem on the behavior of risk for ridge regression models with the optimal regularization constant. The paper is also well written, with very few typos. I congratulate the authors on their polished manuscript. Related work appears to be correctly cited, although this is not my research area so it is hard for me to check. Note that I did not check the theoretical derivations in the appendix so I cannot comment on their correctness.

Weaknesses

The greatest weakness of this paper is that the theoretical results are somewhat incremental and unlikely to have an impact outside of learning theory. In particular, - The connection between ridge regression and sub-sampled ensembles was previously established, so that the main contributions of this work are weakened conditions and new types of equivalences. - It's not clear how interesting it is to answer the conjecture from Nakkiran et al. While the authors prove that the risk for the ridge regression model with optimal regularization constant is monotone increasing in the ratio $d/n$, this only applies to the asymptotic regime and so its practical importance may be limited. - While the authors motivate their work by highlighting connections between ridge regularization and dropout, noisy training, data augmentation, and early stopping. However, these connections are not developed any further and I am skeptical the asymptotic equivalences in this submission will impact those areas.

Questions

Line 93: Is this supposed to mean that Assumption 2 defines RMT features? It's not obvious because the acronym RMT is not defined anywhere and not used in Assumption 2. Line 176: I suggest Changing the name of Assumption 2 from "Feature Vector Distribution" to "RMT" features as well as including a definition for the initialism/acronym, since it isn't stated anywhere. Line 156 and Definition E.1: Does $C_p$ need to be bounded away from zero? Otherwise the asymptotic equivalence definition will be meaningless when $C_p = 0$ almost surely for every $p$. Or perhaps when you say "any sequence $C_p$" you mean "every sequence"? Line 183: Shouldn't this solution be $(\\lambda, \\bar{\\psi})$? Figures 1/2: The paths between equivalent models don't appear to be linear in these figures, although Equation 5 seems to indicate that they are always linear combinations. Is this because the figures are in log-log scale? Theorem 3: I don't understand what is "structural" or "first-order" in this theorem compared to Theorem 1. Is this because the estimators themselves are equivalent, rather than a risk functional of the estimators? While is equivalence of risk functionals "second order"? Theorem 3: "this implies that the predicted values (or even any continuous function applied to them due to the continuous mapping theorem) of any test point will eventually be the same, almost surely, with respect to the training data." The asymptotic equivalence of parameters in Theorem 3 and this statement seem to imply that Theorem 3 covers both Theorem 1 and Proposition 2 by taking $M = \\infty$ --- is this correct? If yes, what is the novelty of these two previous results given Theorem 3? Proposition 4: Is Equation 8 always guaranteed to admit a solution? Moreover, what is the utility of this result when the equivalence only holds asymptotically? That is, when $n \rightarrow \infty$ and the root-finding problem defined by Equation (8) is impractical to solve.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

As mentioned in the "Weaknesses" section, I think the main limitation of this submission is the impracticality of the main theoretical results. Since they extend and generalize a previous result showing asymptotic equivalence of the risk, I do not see the theoretical derivations having an impact on practice. Furthermore, the connections to high-interest topics in ML like early stopping and dropout seem tenuous at best. Since I am not actively involved in learning theory research, I cannot comment on the importance of this work for other members of this community. It would be nice if the authors could provide additional context for their work, including some comments on the novelty of their proof techniques and so on. That way I can better understand the impact on this specific research community.

Reviewer DNS56/10 · confidence 4/52023-07-19

Summary

This work compares the ridgeless ensemble and the ridge estimators in the proportional limit setting (i.e., $d/n\to \phi$). Prior works [11,12,13] show that these two estimators achieve the same out-of-sample risk. The main contribution of this paper has been to (1) weaken the assumptions and (2) broaden the equivalence from out-of-sample risk to other risks such as empirical risk, in-sample risk, and transfer learning risk. Based on this theory, this work also shows that the out-of-sample risk achieved by the optimal ridge estimator monotonically increases as a function of the data aspect ratio (i.e., $\phi$).

Strengths

+ Excellent presentation. I have enjoyed going through the paper. + Compared to [11,12,13], this work has shown equivalence between ridgeless ensemble and ridge estimators in a broader sense and under weaker assumptions. Notably, in this work, the data is allowed to be misspecified and the limiting covariance spectrum needs not to exist, and the proved equivalence holds for transfer learning risk, in-sample risk, empirical risk beside out-of-sample risk. + Section 5 is a neat addition to the main results, showing that the optimal ridge induces an out-of-sample risk increasing with respect to the data aspect ratio.

Weaknesses

- The Conjecture 1 in [1] is stated for every finite $n$ and $p$. Section 5 in this work only applies to the proportional limit setting, where $p/n\to \phi, n,p\to\infty$ for a finite $\phi$. While I think Section 5 is very interesting, I am not sure it is proper to claim that "This resolves a recent open problem raised by Nakkiran et al. [1] under general data distributions and mild regularity conditions." - The theory is limited to the proportional limit regime. To what extent the theory holds in the finite sample regime is unclear. - The technique novelty could be further clarified. - I tried to read the proof and believe they are mostly correct. However, there are a number of typos that confuse me (and prevent me from spending more time checking the proof). For an incomplete list: 1. The equation after Line 514. Missing a factor of 2 in the cross-term. 2. Line 520. $B = AX - I$, missing $-I$. 3. Line 534. $v(- \lambda_1; \psi_1) = v(- \lambda_2; \psi_2)$. $\lambda_2$ instead of $\lambda_1$. 4. The first inequality after Line 622. Left hand side of the inequality, $R(0; \phi, \psi)$ should be $R(0; \phi_1, \psi)$?

Questions

Upto my quick glance, it seems the proof methods are largely built upon existing works such as [12] and [13]. I understand that this work has derived broader equivalence results under weaker assumptions compared to [11,12,13]. Would you please clarify what are the new ingredients in this work that allow so?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

N/A

Reviewer 96Gw2023-08-10

The authors addressed the only minor point I raised. I confirm my initial review and grading, and thank the authors for replying to my curiosity on the possible breaking of the equivalences they prove.

Reviewer DNS52023-08-10

Thank you for your response

Thank you for your response. This paper makes several extensions to prior works on the equivalence between ridgeless ensemble and ridge. However, the theoretical results in this work are still limited to the proportional regime as in prior works. So the contribution of this work is not super high given prior works from my perspective. Also, it is worth noting that this work partially resolves Conjecture 1 in [1] in the proportional limit regime (rather than in the finite-sample/finite-dimensional regime). Therefore, I'd like to maintain my initial review and rating.

Reviewer wEHo2023-08-14

Many thanks for responding to my review and answering my questions. **[W2b]** Right, I agree that this can be interesting. However, it is a bit strong to say "more data can hurt performance" when the results only apply to the limiting ratio of $d/n$. It seems more accurate to say "more data relative to features". I also think it is appropriate to qualify the claim of having resolved this open question as per Reviewer DNS5's comment that this applies to the proportional limit regime only **[Q6]** and **[Q7]** Thanks for clarifying these issues. I now follow why these are first-order or "structural" results. I think it would be useful to remind unfamiliar readers of the type of convergence proved after Theorem 3, i.e. that it is not almost sure convergence of the estimator difference but of linear functional of the difference. **[Q8]** Great. I think this is worth commenting on in the paper. Overall, I think this is a nice submission. The subject area is niche, but the paper is well-executed and the author response has been helpful. I will consider raising my score after the discussion with the other reviewers.

Reviewer nskG2023-08-19

Thank you for the interesting comments and clarifications in the rebuttal. I am maintaining my score. Good luck!

Reviewer MtFr2023-08-21

Thank you for the rebuttal

I am happy with my initial assessment of the paper. Please make the necessary revisions to improve the quality of presentation.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC