Summary
This article presents a refreshing take on losses used in ML, with solid motivational examples and a several of interesting and beautiful results in pure convex analysis.
Strengths
The Fitzpatrick function is a well-known tool for achieving very pure-theoretical results like Minty's theorem in convex analysis. However, before reading this article, I was under the impression that the Fitzpatrick function was just a theoretical tool with limited applications in-practice...
Until now! This paper takes a beautiful tool from pure convex analysis, spruces up some very interesting and new results, and presents it as a new application in ML with some very solid justification concerning sharper inequalities in comparison to the classical class of Fenchel-Young inequalities.
Weaknesses
My main regret is that, due to the reviewing load, I was not able to fully check all of the proofs presented in the appendix. The claims *in the article* are reasonable, interesting, well-presented, and thorough; however, I am quite concerned that there may be some (unintentionally) omitted hypothesis, or holes in the proofs in the appendix (e.g., please also see my statement about Compactness of the domain in "limitations" section), because this often happens in ML conference articles. I wish that an appendix counted as the "reviewed" part of the article, but alas I simply do not have the time to check every detail. I am going to suggest that this paper is accepted, but I suggest that the authors please take care to revise based on the following minor points.
Minor comments:
- Citations appear to be out of numerical order when they are cited
- Line 30: In what sense is "proper" meant? All losses I'm aware of
are proper functions (using the convex analysis definition).
- Line 56: Is R_+ numbers which are strictly positive, or just
positive?
- Line 59: The set must be nonempty for the projection to be defined.
- Line 62: This identity only holds when \Omega is differentiable *and convex*! Otherwise, this definition of the subdifferential produces the emptyset in areas of concavity.
- Line 66: Please provide a citation.
- Line 68: Does this function need to be Legendre?
- Line 101: Monotone is defined here, but mentioned earlier in the
article; would be nice to have the definition at the beginning.
- Definition 1: Within the text of the definition, "dom Omega" appears before Omega is introduced.
- A bit more discussion would be nice on specifically which problems (4) can be calculated.
- Line 132: "Above expression" referring to "\nabla^2 Omega" is ambiguously defined, since in the constrained case Omega is not differentiable.
Questions
- Line 69: As far as I've seen in most convex analysis books and papers, this convention is a bit nonstandard; typically, "+infinity + (-infinity)" is *undefined*. Please comment on which results in this article break if the quantity "+infinity + (-infinity)" is undefined. I.e., if this is "not a valid move", which of your results still hold? It is very important to clarify which results hold under varying algebraic models of convex analysis.
- Line 157 / Prop. 7: the relationship between the claim in the equation following Line 157 and Proposition 7 is unclear. How does the claim in line 157 follow from Proposition 7?
Limitations
The mathematically precise statements (propositions, lemmas, theorems/etc.) should state **all** of their assumptions. In particular, the authors do not repeat that a compact domain of the objective function is required for several of their results. (Unless I'm missing something and compactness is not actually required?)
The authors do a great job of explaining the results of their experiments. More numerics is always a plus, but since the article has so many strong theoretical contributions, I am quite happy with this as-is.