Rebuttal
We thank the reviewer for their feedback. Below, we address their comments.
> The main setup is quite confusing to me. The paper first states that "$f_{\mathsf{weak}} \in \mathbb{R}^d$" is the object we learn. Normally, the model is a function, not a vector, so this was not immediately clear. It is defined later in line 347 how we learn $f$, which is quite far from where it was introduced (line 184). It would be better to define that we train $f$ by MNI earlier.
- We apologize for the confusion with the notation. We have uploaded an updated version where we clarify this, including clarifying how $f$ is trained by MNI earlier. Our goal with the discussion around Line 184 was to encapsulate many different w2s training schemes, but we agree it is helpful to keep a concrete algorithm in mind.
> In line 201, it says, "As a consequence of our main results in Section 3, we will show that the above desiderata are achievable in a simple toy model; see Theorem 3.3 for a formal statement." However, Theorem 3.3 only considers desiderata 1.2 and 2.1, not the entirety of the desiderata.
- The two equations in Lines 421-422 contain the conditions for the additional Desiderata being satisfied. In the revision, we have stated more explicitly that Desiderata 1.i-1.iii are all satisfied, and changed the tags to make it more intuitive. In addition, Remark 3.4 discusses Desiderata 2.i and 2.ii. We felt it would be distracting to focus on the bonus desiderata too much, so we moved it to the Remark. We updated the writing near Line 201 to reference Remark 3.4 regarding the bonus desiderata.
> What is $t$ in Equation (3) of Theorem 3.1?
- $t \in [0, s)$ controls the number of label classes in the multiclass problem: $k = n^t$, following Definition 2. We have updated the wording of the theorem to make this more explicit.
> The notation $u,p,q,r$ used is not very intuitive, and it makes the result difficult to interpret. Is there a simpler way to rephrase the result?
- We apologize for the confusion. The reason we chose this type of notation is that prior work in this area has used similar conventions (see, e.g., [1-4]). We thought it might be easier to work in log space (where conditions are additive) as opposed to multiplicative, but we can try to add a couple sentences to rephrase things.
[1] Wang, K., & Thrampoulidis, C. (2022). Binary classification of gaussian mixtures: Abundance of support vectors, benign overfitting, and regularization. SIAM Journal on Mathematics of Data Science, 4(1), 260-284.
[2] Wang, K., Muthukumar, V., & Thrampoulidis, C. (2021). Benign overfitting in multiclass classification: All roads lead to interpolation. Advances in Neural Information Processing Systems, 34, 24164-24179.
[3] Muthukumar, V., Narang, A., Subramanian, V., Belkin, M., Hsu, D., & Sahai, A. (2021). Classification vs regression in overparameterized regimes: Does the loss function matter?. Journal of Machine Learning Research, 22(222), 1-69.
[4] Wu, D., & Sahai, A. (2024). Precise asymptotic generalization for multiclass classification with overparameterized linear models. Advances in Neural Information Processing Systems, 36.