Summary
This paper proves the necessary and sufficient condition for a loss function to be adversarially consistent.
In the previous literature, either adversarial consistency for restricted hypothesis spaces or negative results for adversarial consistency has been known.
This paper follows this research line to provide a general condition to characterize loss functions.
The condition only requires $C\_\\phi^\*(1/2) < \\phi(0)$, which is quite simple to check.
This even holds for nonconvex loss functions.
The proof technique relies on the strong duality and complementary slackness results between adversarial surrogate loss minimization and the optimal coupling between benign and adversarial distributions.
Strengths
- Very general condition to characterize adversarially consistent losses: Unlike previous results of adversarial $\\mathcal{H}$-consistency (Awasthi et al. (2021)) and negative results for adversarial consistency (Meunier et al. (2022)), this work contributes to show the necessary and sufficient condition for a loss function to be adversarially consistent. This is a new insight into the community and may help design loss functions in adversarial training.
- The condition applies even to nonconvex losses: Traditionally, the theory of calibrated losses mainly concerns convex losses, such as Bartlett et al. (2006) because their proof technique essentially relies on the first-order optimality condition when characterizing loss minimizers. This is a transparent proof technique yet excludes nonconvex losses. In contrast, the proof technique of this paper first translates the optimality of the adversarial loss into the optimal coupling (Propositions 4 and 5) and then deals with the standard consistency analysis for the adversarial distribution $\\mathbb{P}^\*$ (by leveraging Lemma 1).
Weaknesses
Overall, I do not see any concerns about this paper.
There are a few minor comments and questions, which are mentioned in the following "Questions."
Questions
- (Comment) In the introduction, you may consider emphasizing the main result shown in this paper is related to consistency for all measurable functions, not $\\mathcal{H}$-consistency, to clarify how the result differs from the previous works.
- (Question) Regarding Proposition 2: The counterexample $\\mathbb{P}\_0 = \\mathbb{P}\_1$ seems very malicious and rarely happens in practice. Are there any other counterexamples for which the corresponding loss function is not adversarially consistent?
- (Comment) In the definition of $W\_\\infty$, it is better to explain what $(x,y) \\sim \\gamma$ does mean.
- (Typo) In l.233 "fo" -> "of"
- (Typo) In l.260 $R$ -> $R^\\epsilon$
- (Comment) In l.275, it seems better to discuss the existence of the coupling $\\gamma\_i$.
- (Comment) In l.307, I don't think Theorem 4 immediately implies adversarial consistency because it is unclear whether the inequality is tight for any distributions.
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The authors mention the limitations in conclusion: The extension of the convergence rate for general loss functions is left open.
This is theoretical work, and potential negative societal concerns are not applicable.