Exact Generalization Guarantees for (Regularized) Wasserstein Distributionally Robust Models

Wasserstein distributionally robust estimators have emerged as powerful models for prediction and decision-making under uncertainty. These estimators provide attractive generalization guarantees: the robust objective obtained from the training distribution is an exact upper bound on the true risk with high probability. However, existing guarantees either suffer from the curse of dimensionality, are restricted to specific settings, or lead to spurious error terms. In this paper, we show that these generalization guarantees actually hold on general classes of models, do not suffer from the curse of dimensionality, and can even cover distribution shifts at testing. We also prove that these results carry over to the newly-introduced regularized versions of Wasserstein distributionally robust problems.

Paper

References (49)

Scroll for more · 37 remaining

Similar papers

Peer review

Reviewer 6D6k5/10 · confidence 3/52023-07-01

Summary

The paper provides theoretical results characterizing the generalization capabilities of methods based on Wasserstein distributionally robust approaches. In particular, the results presented extend the conditions under which the performance guarantees are not affected by the curse of dimensionality and are applicable for general classes of models.

Strengths

The paper shows that the usage of Wasserstein radius of the order 1/sqrt(n) can provide generalization bounds in situations more general than those considered in existing works (linear models). In addition, the results presented also cover regularized versions of WDRO. The more general results are obtained using a novel type of proof based on a concentration bound for the dual problem, which is of independent interest.

Weaknesses

The paper contribution with respect to the state of the art needs to be better described. In particular, the extension to non-linear models of the scaling 1/sqrt(n). The problematic dimension-dependent scaling arises in Wasserstein methods while other techniques based on robust risk minimization have been shown to provide performance guarantees with the scaling 1/sqrt(n). It would be good if the authors describe this fact and the related work. If I am not mistaken, examples 3.6 and 3.7 correspond to cases for which the right scaling was already proven in previous works. In order to better assess the paper's contribution, it would be good if the authors discuss interesting examples for which the paper provides the right scaling while existing results cannot.

Questions

Would it be possible to include numerical results describing the theoretical results presented?. The choice of the radius in practice is often problematic in Wasserstein methods. The theoretical results provide the scaling of such radius but not a concrete recipe to choose it. It would be useful to explore choices of such radius with the right scaling that result in small error and provide performance guarantees.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The paper adequately describes the limitations of the methods proposed, mostly in terms of the specific assumptions needed for the results to hold.

Reviewer A8VY7/10 · confidence 5/52023-07-05

Summary

This paper presents generalization bound for Wasserstein DRO and entropic regularized Wasserstein DRO (or called Sinkhorn DRO in Wang et al.) formulations. Those generalization bounds do not suffer from the curse of dimensionality. The theoreical analysis is also supported by two examples in Section 3.4.

Strengths

- The theoretical analysis is interesting from two aspects. First, the authors reveal that the radius selection of WDRO to make the empirical robust loss dominate the true loss does not suffer from the curse of dimensionality. The analysis follows different techniques from existing literature such as Gao et al, Blanchet et. al, etc. Second, the technique is general enough so that it also applies to entropic regularized Wasserstein DRO (or called Sinkhorn DRO in Wang et al.) formulations. This is the first work that investigates the statistical properties of such formulations. - The authors also present two examples in machine learning to demonstrate the technique assumption holds and the proposed theoretical analysis applies.

Weaknesses

- The writing of this paper could be potentially improved: 1. There should be a comma in Eq.(1), or equation between line 83-84, or Eq. (4), or equation between line 194-195, or equation between line 301-302. 2. There should be a period in Eq. (10). 3. The contribution and related work part in the introduction section should be separated. 4. It would be a little bit confusing to first introduce KL-divergence regularized WDRO risk in Eq.(5-6) and then introduce it corresponds to the Sinkhorn ambiguity set in line 194-195. The authors should put them together in Section 2.2 5. The notation could be potentially improved. For example, in Eq. (7) the authors use $\hat{\mathcal{R}}$ to refer to the risk based on empirical distribution $P_n$. I would suggest replace the notation $P_n$ with $\hat{P}_n$ for consistency. Further, in Eq. (7) I think the authors are meaning $\rho$ should at least scale in the order of $\sqrt{(1+\log(1/\delta))/n}$, then why not write $\Omega(\sqrt{(1+\log(1/\delta))/n})\le \rho$ instead of $O(\sqrt{(1+\log(1/\delta))/n})\le \rho$? The same applies for equation between line 199-200. - It is great that the authors present statistical analysis for entropic regularized Wasserstein DRO. I would suggest the authors add some explanation or numerical example to demonstrate the benefit of introducing entropic regularization. Will it bring extra benefits than standard WDRO? - The analysis is limited to quadratic cost function, which could be restrictive. From my own trial and reading, I think the major difficulty for generalization is that, it is difficult to apply Laplace approximation technique for general p-th power of norm function. In other words, it is difficult to obtain the p-th power of norm counterpart of Lemma A.3 and Lemma G.1.If so, I suggest the authors add explanation for the difficulty of extension. - Some literature is missing. For example, readers may wonder why consider adding entropic regularization to WDRO problem and what is the applications? I suggest the authors make the following revisions: 1. update reference [J. Wang, R. Gao, and Y. Xie. Sinkhorn distributionally robust optimization. arXiv preprint arXiv:2109.11926, 2021] as [J. Wang, R. Gao, and Y. Xie. Sinkhorn distributionally robust optimization. arXiv preprint arXiv:2109.11926, 2023]. In the updated version, the authors demonstrate that people can find $\delta$-optimal solution to general entropic regularization WDRO problem with complexity $\tilde{O}(1/\delta^2)$. So one major benefit of adding entropic regularization is the computational tractability; 2. add several application papers brought by entropic regularization WDRO in literature review: (i) Dapogny, Charles, et al. "Entropy-regularized Wasserstein distributionally robust shape and topology optimization." Structural and Multidisciplinary Optimization 66.3 (2023): 42. (ii) Song, Jun, et al. "Provably Convergent Policy Optimization via Metric-aware Trust Region Methods." arXiv preprint arXiv:2306.14133 (2023). (iii) Wang, Jie, and Yao Xie. "A data-driven approach to robust hypothesis testing using sinkhorn uncertainty sets." 2022 IEEE International Symposium on Information Theory (ISIT). IEEE, 2022. (iv) Wang, Jie, et al. "Improving sepsis prediction model generalization with optimal transport." Machine Learning for Health. PMLR, 2022.

Questions

N/A

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer wf5D7/10 · confidence 3/52023-07-05

Summary

This work proves generalization guarantees for Wasserstein DRO models that only require the radius of order $O(n^{-1/2})$ under mild assumptions for general classes of models. This provides concentration results that do not suffer from the curse of dimensionality.

Strengths

The theoretical contribution is the main strength. The empirical concentration of Wasserstein distance suffers from the curse of dimensionality, and this paper is able to prove the results (under some assumptions) that do not have this curse of dimensionality issue and provide statistical guarantees on the performance of WDRO solutions.

Weaknesses

There is no significant weakness in this paper. Nevertheless, I think adding some discussion or examples for which the assumptions and thus the results in this paper do not hold can be beneficial; it can show failure cases and may also motivate future directions for the extension.

Questions

It might be better to split Section 3 into two shorter sections for better readability. And the discussion may also be extended by adding examples of failure cases with potential methods of relaxation.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

yes

Reviewer ZTJQ7/10 · confidence 2/52023-07-05

Summary

This paper provides generalization guarantees of Wasserstein DRO for a general class of functions, in which the radius scales as $1/\sqrt{n}$ and does not suffer from the curse of dimensionality. Moreover, these guarantees hold for any distribution in the neigbourhood of the true distribution, so that they still apply when the distribution shifts at testing time. The results in this paper hold for both constrained and regularized version of Wasserstein DRO. The authors also provide a proof sketch that explains the main ideas and techniques used in the proof, and apply their results to logistic and linear regression.

Strengths

1. This paper provides novel generalization guarantees such that the robustness radius does not suffer from the curse of dimensionality. To the best of my knowledge, the results in this paper are novel and make a non-trivial contribution to the DRO community. Moreover, the authors consider the regularized version of Wasserstein DRO and provide similar guarantees as well. 2. Most parts of the paper are well-written. The necessary backgrounds are clearly explained, and theorems are accompanied with detailed explanations of related definitions and concepts.

Weaknesses

In Section 3.4, the authors considers logistic and linear regression as applications of their theorems. It would be better if more complicated and popular parametric models can be included in this section to justify the main assumptions.

Questions

1. How does the setting considered in this paper compared with other works? Could you give a brief and high-level discussion of why the generalization guarantee is dimension-independent in your setting? 2. Is it possible to obtain similar generalization guarantee for DRO with $\phi$-divergence?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

This paper does not have potential negative societal impact.

Reviewer A8VY2023-08-10

After reading the rebuttal

I have read the rebuttal and I am happy to raise my score to 7.

Reviewer 6D6k2023-08-15

I thank the authors for their responses that mostly address my comments/questions. I believe the paper deserves to be published and it will be improved in the camera ready version

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC