Precise Asymptotic Generalization for Multiclass Classification with Overparameterized Linear Models

We study the asymptotic generalization of an overparameterized linear model for multiclass classification under the Gaussian covariates bi-level model introduced in Subramanian et al.~'22, where the number of data points, features, and classes all grow together. We fully resolve the conjecture posed in Subramanian et al.~'22, matching the predicted regimes for generalization. Furthermore, our new lower bounds are akin to an information-theoretic strong converse: they establish that the misclassification rate goes to 0 or 1 asymptotically. One surprising consequence of our tight results is that the min-norm interpolating classifier can be asymptotically suboptimal relative to noninterpolating classifiers in the regime where the min-norm interpolating regressor is known to be optimal. The key to our tight analysis is a new variant of the Hanson-Wright inequality which is broadly useful for multiclass problems with sparse labels. As an application, we show that the same type of analysis can be used to analyze the related multilabel classification problem under the same bi-level ensemble.

Paper

References (65)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer qo1L7/10 · confidence 1/52023-07-04

Summary

The paper studies the asymptotic generalization error behavior of an overparameterized linear model and under the Gaussian covariates bi-level model. In this setup, the number of data points, features, and classes all grow together. The authors manage to fully characterize the regimes of the generalization error, which is surprisingly “polarized”. This solves the conjecture posed by Subramanian et al. (2022).

Strengths

Both the achieved theoretical result and also the used technical tools (like the newly established Hanson-Wright inequality) are highly novel and strong. Even though I am not familiar with the literature, I estimate that the work is quite valuable.

Weaknesses

- The considered assumptions for the distribution of features are very simplistic. Both independence and having identical Gaussian distributions are very restrictive; which highly influences the practicality of the results. - The paper studies the generalization error of linear models; which are far from the deep neural networks. In this sense, there is still a big gap in the practical aspects of the paper (and also the previous literature on this). However, it is totally understandable that these are the first steps toward that goal. - The paper lacks providing the needed intuitions about the obtained results.

Questions

- It would be useful to provide more intuition about the considered “bi-level ensemble” model. In this model, it seems that while we are considering $n^p$ dimensions, the ``effective dimensions’’ is the favored ones; meaning that the features in the rest of the dimensions (and their relative magnitude) somehow either add a “useful noise injection” or become dominant with respect to the “signal” which makes the prediction impossible. This might be inexact, but this is just to give an example of what kind of intuition I refer to. - Similarly, it would be extremely useful to discuss the different regimes of the theorems. Why having $r$ close to 1 would fail the learner? Why having a small $p-(q+r)$ would do so? - Finally, the scope of the paper remains a bit narrow (and purely technical) and the authors could probably use the intuitions derived from their results to discuss some potential understanding that these results could give about some more realistic setups. This is shortly done in the discussion section, but it would be appreciated to be extended.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

4 excellent

Limitations

As discussed above, and I guess the authors would also agree, the considered model is quite restrictive.

Reviewer V1nE8/10 · confidence 3/52023-07-06

Summary

The paper titled "Asymptotic Generalization of Overparameterized Linear Models for Multiclass Classification under Gaussian Covariates Bi-level Model" presents a study on the asymptotic generalization of overparameterized linear models for multiclass classification under the Gaussian covariates bi-level model. The authors provide an asymptotic characterization of the generalization of a linear model for multiclass classification in an idealized Gaussian setting where a) the number of data points, b) the dimension, and c) the number of classes diverge while their ratio remains fixed. An interesting result, in particular, is that the min-norm interpolating classifier can be suboptimal in this regime.

Strengths

The paper present many strong analytical results. In particular, the authors have successfully resolved a conjecture posed in a previous works (Subramanian et al. '22,) and established new lower bounds that demonstrate the misclassification rate either goes to 0 or 1 asymptotically. The paper also introduces a new variant of the Hanson-Wright inequality, a tool used in high-dimensional probability, that is particularly useful for multiclass problems with sparse labels. Sparse labels refer to situations where only a small number of classes are represented in the dataset.

Questions

How realistic is the Bi-level feature weighting model ? I am thinking instead of the source/capacity setting where I would expect a more power-law-like behavior. How different would the conclusions be?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The paper is theoretical in nature and its limitations are stated in the theorems

Reviewer S1af8/10 · confidence 1/52023-07-09

Summary

This is a theoretical paper that gives insight into how and when an overparametrized linear classification model, for multi-class classification, can be successfully generalized. In particular, they look at multiclass classification under the Gaussian covariates bi-level model introduced by Subramamian et al. in 2022, and fully resolve a conjecture from that paper. The key to their analysis is a new variant of the Hanson-Wright inequality.

Strengths

This paper builds on previous work to give tight bounds on the regions where generalization is possible, improving significantly on previous partial results. It also makes rigorous previous analyses based on heuristic calculations. This is a very well-written paper. Based on the proof sketch given in the main paper, the technical level of the proofs seems high, involving both the use of existing techniques as well as the proof of a new version of the Hanson-Wright inequality,.

Weaknesses

No significant weaknesses were noted.

Questions

None.

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

It would be helpful to comment on the limitations of the bi-level model studied.

Reviewer ggP88/10 · confidence 4/52023-07-21

Summary

In their main result, Theorem 3.2, the authors establish Conjecture 3.1 which is a conjecture posed in 2022 describing the asymptotic misclassification probability of the bi-level ensemble model (Definition 1) under a sparsity assumption (Assumption 1). They provide a rigorous and tight analysis, with very clear explanations both in-text and in the appendix.

Strengths

The paper is clearly written and explained wonderfully, the proofs are detailed, and the contributions are interesting. The proofs are sound and clearly articulated. Minor point, the appendix is also very neatly organized which makes proofreading nice. **Update:** My questions and concerns have been addressed.

Weaknesses

Overall the paper is rigerous and clear, so I do not really have any significant weakness worth reporting; only a few minor comments. * Only minor comments:* - In equations (36)-(40) could you explicitly write the polylogarithmic terms in the denominator, or at the very least, precisely defined them after the equation environments. - Could you formally state that all r.v. are defined on the same probability space, at the beginning of the paper; for rigor. - Very minor point, above equation (219), I guess $\boldsymbol{A}_{-\boldsymbol{S}}^{-1}$ is *block diagonal* not diagonal.

Questions

There are two little details, which I didn't fully follow, so let me ask: - Perhaps a silly question, but in equation (182), are you using a Brascamp–Lieb inequality (or something simpler which I'm possibly missing)? - In equation (240) do you mean $n^{t-1}\hat{\boldsymbol{f}}_1[1]$ is $\Theta(n^{(p-q-2)/2})$?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

N/A

Reviewer PSb67/10 · confidence 2/52023-07-27

Summary

In this paper, the authors analyze the generalization of the linear multiclass classifiers in the overparametrized regime under the bi-level model with Gaussian covariates. In particular they prove a conjecture made in a previous paper about the region (characterized by the paramters of the bi-level model) under which the generalization error can go to zero. In fact, they establish a `0-1` law for generalization error: depending upon the parameter regime the probability of error either converges to zero or one. In the process of proving this result the authors also establish a generalization of the usual Hanson Wight inequality for quadratic forms to exploit "soft sparsity". Overall, I feel that this paper contains a technically rigorous analysis of the overparametrized classification problem somewhat restrictive model assumptions.

Strengths

The paper is well written, and while the material is quite technical, the authors do make an effort to provide sufficient context and explanations to motivate them. In particular, the detailed overview of the argument used in proving the main result (Theorem 3.2) is quite helpful in getting the idea used in the proof.

Weaknesses

1. While this paper does not seem to have any obvious flaws, I feel that the model analyzed in this paper is a bit too restrictive. In a footnote on page 4, the authors mention that "such models are widely used to study learning even beyond this particular thread of work". It would be very helpful, if the authors include a detailed discussion of at least one such practical application, which is naturally modeled by the bi-level model studied in this paper. 2. Besides the model, even the construction of the classifier makes some strong assumptions. In particular, the classifier is constructed after reweighting the features with weights $(\lambda_i)_{i \geq 1}$ that are tuned to the specific bi-level model. Since in most practical problems, there is at least some level of misspecification, I feel that this reduces the practical utility of the results in this paper.

Questions

1. Can you include a discussion of some practical tasks where the bi-level model arises naturally? 2. Regarding the second point in the weaknesses section, can you discuss (at least informally) the effect of model misspecification on the generalization guarantees. More generally, do you think that the same performance guarantees as Theorem 3.2 (i.e., probability of error converging to zero for the same parameter ranges) can be achieved adaptively without using the knowledge of model parameters ($p, q, r, t$) to construct the classifier? Or there is a price to pay for achieving zero generalization adaptively?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

I think that the strong model assumptions used in this paper might reduce the practical utility of the results of this paper.

Reviewer PSb62023-08-18

Reply to the rebuttal

Thank you for the clarifications about the bi-level model and model-misspecification. I am happy to update my score to 7.

Reviewer BQHA6/10 · confidence 1/52023-07-27

Summary

By resolving the conjecture posed by Subramanian et al. (2022), the authors address the asymptotic generalization for overparameterized minimum-norm interpolation (MNI) linear multi-class classifiers under two assumptions: (1) features are Gaussian vectors and labels are generated from $1$-sparse noiseless model; (2) the scaling follows bi-level ensemble. This paper can be seen as a completion to the previous result of Subramanian et al. (2022) since it captures the generalization for both the regime where the regressor fails and regime where it works. The technical contribution appears to be a new variant of the Hanson-Wright inequality.

Strengths

A sound completion to the previous theoretical result.

Weaknesses

The setting of bi-level ensemble seems a bit weird to me, especially the eq (12). Maybe the authors can offer more explainations.

Questions

Line 165: "it suffices for the maximum value of the LHS ......'' Is it the maximum instead of minimum here?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors adequately addressed the limitations.

Reviewer ggP82023-08-16

Happy with edits

Dear authors, Thanks you very much for the clear response and very nice paper. Goodluck :)

Reviewer BQHA2023-08-16

Dear authors, Thank you for the clarifications on eq(12) and it makes more sense to me now.

Area Chair ZMoe2023-08-21

To the authors: your response has been read and is being considered.

Area Chair ZMoe2023-08-21

To the authors: your response has been read and is being considered.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC