Summary
In this paper, the authors analyze the generalization of the linear multiclass classifiers in the overparametrized regime under the bi-level model with Gaussian covariates. In particular they prove a conjecture made in a previous paper about the region (characterized by the paramters of the bi-level model) under which the generalization error can go to zero. In fact, they establish a `0-1` law for generalization error: depending upon the parameter regime the probability of error either converges to zero or one. In the process of proving this result the authors also establish a generalization of the usual Hanson Wight inequality for quadratic forms to exploit "soft sparsity".
Overall, I feel that this paper contains a technically rigorous analysis of the overparametrized classification problem somewhat restrictive model assumptions.
Strengths
The paper is well written, and while the material is quite technical, the authors do make an effort to provide sufficient context and explanations to motivate them. In particular, the detailed overview of the argument used in proving the main result (Theorem 3.2) is quite helpful in getting the idea used in the proof.
Weaknesses
1. While this paper does not seem to have any obvious flaws, I feel that the model analyzed in this paper is a bit too restrictive. In a footnote on page 4, the authors mention that "such models are widely used to study learning even beyond this particular thread of work". It would be very helpful, if the authors include a detailed discussion of at least one such practical application, which is naturally modeled by the bi-level model studied in this paper.
2. Besides the model, even the construction of the classifier makes some strong assumptions. In particular, the classifier is constructed after reweighting the features with weights $(\lambda_i)_{i \geq 1}$ that are tuned to the specific bi-level model. Since in most practical problems, there is at least some level of misspecification, I feel that this reduces the practical utility of the results in this paper.
Questions
1. Can you include a discussion of some practical tasks where the bi-level model arises naturally?
2. Regarding the second point in the weaknesses section, can you discuss (at least informally) the effect of model misspecification on the generalization guarantees. More generally, do you think that the same performance guarantees as Theorem 3.2 (i.e., probability of error converging to zero for the same parameter ranges) can be achieved adaptively without using the knowledge of model parameters ($p, q, r, t$) to construct the classifier? Or there is a price to pay for achieving zero generalization adaptively?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
I think that the strong model assumptions used in this paper might reduce the practical utility of the results of this paper.