Summary
This paper studies the classic problem of regression under noisy observations, with a convex and closed function class $\mathcal{F}$. In particular, the paper considers the performance of the Empirical Risk Minimizer (ERM) under the mean squared error, in both the fixed design and random design settings.
At a high level, the author(s) show that:
1. The variance of the ERM is comparable to the minimax rate of regression, implying that when ERM is not minimax-optimal, it must be due to large bias
2. ERM is admissible up to constant factors, via a simpler proof using fixed point theorems
3. ERM enjoys stability, in the sense that all almost-empirical-loss-minimizers also have similar expected loss as the actual empirical loss minimizer.
4. For "non-Donsker" function classes, the converse of 3 is false, that there is always some ground truth function $f^* \in \mathcal{F}$ such that, with high probability over the samples and noise, there is a function $f_{\mathrm{bad}}$ whose expected loss close to the ERM, but that its empirical loss is much larger than the ERM's.
I did the following review as an emergency review, so I did not check the details/appendices carefully.
Strengths
The ERM is perhaps the most commonly used regression estimator. This paper furthers our understanding of its behavior, particularly for function classes where ERM isn't minimax-optimal. I found the introduction well-written, describing each result at a high level. I also find the observation that "non-minimax-optimality must be due to bias" to be an interesting result.
Weaknesses
While I liked the paper, and recommend a weak accept, I think there are the following writing/presentational issues that can be improved. Overall, my comment is that the paper currently perhaps reads better for experts who have worked on ERM (or at least, in empirical process theory), but not that easy to read for a more general learning theory audience.
- The paper reads a bit like a collection of related results, but without a "main point", and could maybe be strung together better as a story. For example, the "variance <= minimax-rate" result and the stability result seem somewhat disjoint in the introduction, even though later on in Theorem 1, they are shown together as one big result. I was also a bit confused about the writing in Section 2, in terms of the narrative. For example, I don't quite understand how Theorem 2 "complements" Theorem 1 (cf. Line 150), though Theorem 2 is a cool result itself with the local minimax optimality and is used to prove Corollary 1 (the admissibility result). Moreover, Theorem 3 also "complements" Theorem 1 (cf. Line 185), but there's a quadratic gap. Is the gap necessary?
- I also found that the technical Section 2 is a bit too heavy on definitions and assumptions and not enough interpretation. When I was reading, I felt that the author(s) tried hard to succinctly give the most general results that are shown, but in my opinion, for readability, it might be better to start with simpler cases (and perhaps even informal versions) of statements (particularly the assumptions), and provide more interpretation.
- Related to the above point, some of the assumptions/quantities in the results are stated without much interpretation or intuition (in the main body). There are also a few comments along the lines of "this assumption is considered mild in the literature/by some other authors" without much additional interpretation. This makes the paper not quite as self-contained as it could be. This issue is particularly prevalent in Section 2.2 (especially the isoperimetry assumptions), and it became quite hard to interpret the results. I can see that there are some remarks in the appendices, but I think a lot of them really should be in the main body for readability.
- The term non-Donsker was never defined in the main body, even though it is a key element in a main result. While Donsker classes are a basic object in empirical process theory, I don't think it should be assumed knowledge for the general theory reader. There is also a lack of discussion on whether Theorem 4 only hold for non-Donsker classes, or more generally what's the significance of the assumption: whether it's a necessary condition for the result, or just that this is the result that can be proved.
Questions
Minor questions and comments:
1. Please consider using the same number-counter across all of definitions/theorems/assumptions. It was hard to scroll through the paper looking at backward references when reading.
2. (Line 164) The author(s) mention that $\mathcal{G}_\ast$ in Theorem 2 can replace 2 with any other absolute constant. How does the replacement propagate to Theorem 2? Does it change the constants in the $\asymp$?
3. Theorem 6, the assumptions 1,2,4,6 have nothing to do with the instantiation of $\mathbf{X}$?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.