Thank you very much for your positive feedback, all your comments/remarks, and your two main questions above. It helps us improving our presentation, and especially on the position w.r.t Azizian et al. 2023a, which is the closest work.
In the revision, we now provide a discussion, just below our main assumptions (Assumption 2.1) where we compare our setting to Azizian et al. 2023a in details. This complements the key differences, already mentioned in the related work section (lines 120 to 130).
**Let us recall here our improvements, compared to Azizian et al (2023a):**
- The setting of Azizian et al (2023a) restricts to smooth functions $f \in \mathcal{F}$ (twice differentiable with uniformly bounded derivatives on a convex sample space). We only require the $f \in \mathcal{F}$ to be continuous on a metric space. In addition to nonsmooth functions, this allows us to consider distributions on sample spaces with discrete and continuous variables (as for e.g. classification tasks).
- Their proof require to take $c$ as the squared norm and the
reference distribution $\pi_0(\cdot|\xi)$ as a Gaussian distribution. We consider instead general costs $c$, continuous with respect to a distance on $\Xi$ and an arbitrary reference probability distribution.
For instance, this setting is captured by us but not by Azizian et al. (2023a):
> (i) The sample space $\Xi = B(0,R) \times \{0,1\}$ where $R > 0$
>
> (ii) The loss family $\mathcal{F} = \\{ f(\theta, \cdot) \ : \ \theta \in \Theta \\}$ with the cross entropy loss $$f(\theta, x,y) = - y \log(h(\theta,x)) - (1 - y) \log(h(\theta,x))$$
where $h(\theta, \cdot)$ is a feedforward network with RELU activations and $\Theta$ is compact.
>
> (iii) The cost function: $c((x,y), (x',y')) = \|x-x'\|_{p}^{q} + \kappa \mathbb{1}\_{y \neq y'}$
- Moreover, in their proof, to overcome nonsmoothness of WDRO (which poses concerns for applying concentration results), they require two technical assumptions: a compactness condition (1) and growth conditions around maximizers (2), this is their **Assumption 5**:
(1) For any $R > 0$, there exists $\Delta > 0$ such that
$$\forall f \in \mathcal{F}, \; \forall \zeta \in \Xi, \; d\left(\zeta, \operatorname{argmax} f\right) \geq R \implies f(\zeta) - \max f \leq -\Delta.$$
(2) There exist $\mu > 0$ and $L > 0$ such that, for all $f \in \mathcal{F}$, $\xi \in \Xi$ and $\xi^*$ a projection of $\xi$ on $\operatorname{argmax} f$,
$$f(\xi^*) \geq f(\xi) + \frac{\mu}{2} \|\xi - \xi^*\|^2 - \frac{L}{6} \|\xi - \xi^*\|^3.$$
We do not rely on these conditions. They are rather strong and difficult to verify since maximizers over the sample space are hard to control in general and both depend on the sample space and the function class geometries. In particular, (1) requires $\mathcal{F}$ to be compact with respect to a distance defined by summing the sup norm and the Hausdorff distance between maximizers sets. Equivalently, this can be seen as the continuity of $f \mapsto \operatorname{argmax} f$ on $\mathcal{F}$ (see our Proposition F.4 in appendix), which is hard to verify.
The main technical difficulties in our proof was thus to get rid Assumption 5 from Azizian et al 2023a and to deal with the nonsmooth aspect of the WDRO objective. To this purpose, we simplified the proof and use nice nonsmooth analysis tools. We highlight this aspect in the sketch of proof, section 4.2 ''Definition of the lower bound" where we present a maximal radius function.
**About your smaller question:** Indeed, the assumption is vacuous in this case and we may take $\omega = 1$. We decided to keep the constant $\omega$ in the results in order to highlight the dependence of $\lambda_{\text{low}}$ on the hypothesis domain. To make our statements more precise, we added the condition $\omega \geq 1$ in Assumption 3.1.1.