Strengths
The two contributions presented in this paper are significant because Contribution i) improves on the literature, in particular with respect to the closest work (Kou et al., 2023), which proposes exactly the same analyses but for a gradient descent training; Contribution ii) is, up to my knowledge, the first such result for SAM. The latter result provides a comprehension on the well-known fact that SAM helps networks to generalize better than SGD. This is a real breakthrough, which should be welcomed by the community.
Weaknesses
I was very enthusiastic by the results presented in this paper but I became disillusioned when reading it. The presentation is disordered and in particular, notations are not all defined, making the results (even informally stated) very difficult to understand. After having a glance at the appendix (which seems to be rigorously written), it seems that the manuscript is an assembly of results extracted from the appendix, unfortunately awkwardly built. As for me, it is a pity because I think that the results are important and of interest for the community, but the presentation is inadequate and not clear enough.
Comments:
1) The abstract misses to state Contribution i) (which constitutes a large part of the manuscript and which is clearly reminded in the conclusion).
2) Definition 2.1 is difficult to understand. In particular, the parts “signal contained in each data point” and “the other are given by” are unclear. Besides, even though this model seems common in the recent literature, a discussion regarding its limitations, for instance independence between labels and covariates, would be appreciated. The same remark can be done regarding the architecture of the neural network considered (Section 2.2). In addition, Condition 3.1 is discussed only quickly Line 142 ; a deeper discussion regarding the role of the variances and the polynomial degrees appearing would be appreciated.
3) There is no link between statements/results in the paper and their counterparts in the appendix. Since the reader is supposed to juggle the two, it would be helpful to know where are the formal statements and the proofs in the appendix.
4) I understand that authors try to explain the derivation of their main results but I am not very comfortable with informal statements, first because of the notation problem previously stated and also because some conclusions are given with vague explanations (in my case they are difficult to grasp) while they seem to be quite difficult to obtain formally. This is the case for instance Line 218 regarding the symmetry of $\rho$.
5) Line 117, the authors invoke the discontinuity of the gradient to justify the need of an analysis based on an other technique than the Taylor expansion. As for me an even better argument is that the ReLU function is formally not differentiable everywhere.
6) The statement of Theorem 4.1 is a bit unclear: the parts “we train/we can train” could be replaced by passive forms, “neural networks” refers, as far as I understand, to the chosen architecture. In addition, mentioning SGD in a result concerning SAM may throw the reader. It could be specified beforehand that SAM training uses intrinsically SGD (but at a point which is not the current iterate).
7) Lemma 4.3 is true only for a particular choice of $\tau$, as stated in Theorem 4.1 (I am not sure that this is clear in Lemma C.5). Since this result is important, this should appear in Lemma 4.3 and discussed after.
8) The related work section (Section 6) appears at the end of the manuscript but is an enumeration of papers. Such an enumeration is generally well placed just after the introduction. Placing the related work section after the result statements makes sense if it discusses technical differences with the closest papers. As it happens, this work seems to be based to a certain extent on (Kou et al., 2023). Thus, it would be enriching to discuss the contribution, and particularly the technical novelties (GD → SGD/SAM), of this work with respect to (Kou et al., 2023).
9) Mathematical remarks: Line 47, $t$, $\mathbf W^{(t)}$, $\boldsymbol \mu$, $L_{\mathcal D}^{0-1}$, $p$, $\Omega$ are not defined; Lines 47, 150, 240 and after, the expression “converges to $\epsilon$” should be replaced by “converges to $0$”; Line 66, $l_2$ → $\ell_2$, Line 68, absolute values in $|a_k/b_k|$ seem useless; “omit logarithmic terms” should be defined explicitly; Line 97, $\mathbf W$ is not defined; Line 100, “is a collection of” means “is a matrix”, doesn’t it? Line 101, $[n]$ is not defined; Line 113, it is specified that $\sigma_0^2$ is the variance of the normal distribution but not Line 80; in Equation (5), the 2-norm should be a Frobienus norm; what is the utility of Equation (6) with respect to the equation Line 175? Line 121, $S_i^{(t, b)}$ → $\tilde S_i^{(t, b)}$; Line 219, $T^*$ is not defined; Lines 218 and 226, the authors could remind what are $H$ and $B$; Line 226, I understand that $|\cdot|$ is the cardinality but this is not stated.
10) Typographical remarks: Line 9, “for the certain” → “for a certain”; Line 27, “with minimal gradient” → “with minimal gradient norm”; Figure 1, blue and yellow are inverted; Line 102, “cross-entropy loss function” and “logistic loss” are redundant; Line 115, “indicator function” → “the indicator function”; Line 156, “Bayesian optimal risk” → “Bayes risk”; full points are missing Lines 202, 207, 214; Line 237: “iteration” → “iterations”; Line 251, full point instead of colon; Figure 2, y-label is cut; Line 299, “an generalization” → “a generalization”.