Summary
This paper presents generalization bound for Wasserstein DRO and entropic regularized Wasserstein DRO (or called Sinkhorn DRO in Wang et al.) formulations. Those generalization bounds do not suffer from the curse of dimensionality. The theoreical analysis is also supported by two examples in Section 3.4.
Strengths
- The theoretical analysis is interesting from two aspects. First, the authors reveal that the radius selection of WDRO to make the empirical robust loss dominate the true loss does not suffer from the curse of dimensionality. The analysis follows different techniques from existing literature such as Gao et al, Blanchet et. al, etc. Second, the technique is general enough so that it also applies to entropic regularized Wasserstein DRO (or called Sinkhorn DRO in Wang et al.) formulations. This is the first work that investigates the statistical properties of such formulations.
- The authors also present two examples in machine learning to demonstrate the technique assumption holds and the proposed theoretical analysis applies.
Weaknesses
- The writing of this paper could be potentially improved:
1. There should be a comma in Eq.(1), or equation between line 83-84, or Eq. (4), or equation between line 194-195, or equation between line 301-302.
2. There should be a period in Eq. (10).
3. The contribution and related work part in the introduction section should be separated.
4. It would be a little bit confusing to first introduce KL-divergence regularized WDRO risk in Eq.(5-6) and then introduce it corresponds to the Sinkhorn ambiguity set in line 194-195. The authors should put them together in Section 2.2
5. The notation could be potentially improved. For example, in Eq. (7) the authors use $\hat{\mathcal{R}}$ to refer to the risk based on empirical distribution $P_n$. I would suggest replace the notation $P_n$ with $\hat{P}_n$ for consistency. Further, in Eq. (7) I think the authors are meaning $\rho$ should at least scale in the order of $\sqrt{(1+\log(1/\delta))/n}$, then why not write $\Omega(\sqrt{(1+\log(1/\delta))/n})\le \rho$ instead of $O(\sqrt{(1+\log(1/\delta))/n})\le \rho$? The same applies for equation between line 199-200.
- It is great that the authors present statistical analysis for entropic regularized Wasserstein DRO. I would suggest the authors add some explanation or numerical example to demonstrate the benefit of introducing entropic regularization. Will it bring extra benefits than standard WDRO?
- The analysis is limited to quadratic cost function, which could be restrictive. From my own trial and reading, I think the major difficulty for generalization is that, it is difficult to apply Laplace approximation technique for general p-th power of norm function. In other words, it is difficult to obtain the p-th power of norm counterpart of Lemma A.3 and Lemma G.1.If so, I suggest the authors add explanation for the difficulty of extension.
- Some literature is missing. For example, readers may wonder why consider adding entropic regularization to WDRO problem and what is the applications? I suggest the authors make the following revisions:
1. update reference [J. Wang, R. Gao, and Y. Xie. Sinkhorn distributionally robust optimization. arXiv preprint arXiv:2109.11926, 2021] as [J. Wang, R. Gao, and Y. Xie. Sinkhorn distributionally robust optimization. arXiv preprint arXiv:2109.11926, 2023]. In the updated version, the authors demonstrate that people can find $\delta$-optimal solution to general entropic regularization WDRO problem with complexity $\tilde{O}(1/\delta^2)$. So one major benefit of adding entropic regularization is the computational tractability;
2. add several application papers brought by entropic regularization WDRO in literature review:
(i) Dapogny, Charles, et al. "Entropy-regularized Wasserstein distributionally robust shape and topology optimization." Structural and Multidisciplinary Optimization 66.3 (2023): 42.
(ii) Song, Jun, et al. "Provably Convergent Policy Optimization via Metric-aware Trust Region Methods." arXiv preprint arXiv:2306.14133 (2023).
(iii) Wang, Jie, and Yao Xie. "A data-driven approach to robust hypothesis testing using sinkhorn uncertainty sets." 2022 IEEE International Symposium on Information Theory (ISIT). IEEE, 2022.
(iv) Wang, Jie, et al. "Improving sepsis prediction model generalization with optimal transport." Machine Learning for Health. PMLR, 2022.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.