Bayesian Nonparametrics Meets Data-Driven Distributionally Robust Optimization

Training machine learning and statistical models often involves optimizing a data-driven risk criterion. The risk is usually computed with respect to the empirical data distribution, but this may result in poor and unstable out-of-sample performance due to distributional uncertainty. In the spirit of distributionally robust optimization, we propose a novel robust criterion by combining insights from Bayesian nonparametric (i.e., Dirichlet process) theory and a recent decision-theoretic model of smooth ambiguity-averse preferences. First, we highlight novel connections with standard regularized empirical risk minimization techniques, among which Ridge and LASSO regressions. Then, we theoretically demonstrate the existence of favorable finite-sample and asymptotic statistical guarantees on the performance of the robust optimization procedure. For practical implementation, we propose and study tractable approximations of the criterion based on well-known Dirichlet process representations. We also show that the smoothness of the criterion naturally leads to standard gradient-based numerical optimization. Finally, we provide insights into the workings of our method by applying it to a variety of tasks based on simulated and real datasets.

Paper

Similar papers

Peer review

Reviewer p57R6/10 · confidence 3/52024-07-04

Summary

This article proposes a new method for optimisation of risk under uncertainty, using Dirichlet Processes to introduce an extra degree of robustness to modelling uncertainty on the data generating process. The method is generally applicable to many statistical learning problems through the use of loss functions, and further links are made to the economic decision-making literature. The theoretical properties of the methods are explored, and a Monte Carlo approximation is introduced to make tractable the inference for the DP model. Experiments are performed on three simulated experiments, and three real data sets.

Strengths

This article represents a novel contribution to a very general problem, with the use of a neat mathematical representation of a statistical decision maker’s uncertainty in the form of a Dirichlet Process. The relevant theory behind the method is explored thoroughly, and there are several different applications explored. The quality of the scientific writing is fairly high. The links made with the economic decision making literature are interesting, helping elucidate the underlying point about ambiguity aversion.

Weaknesses

The idea of placing a prior on the data generating process instead of the parameters is unusual from the traditional Bayesian perspective. While this is not necessarily wrong, the authors do not do a particularly good job of communicating the implications of constructing a prior in this way, or how existing intuitions that the readership may have can be transferred to this new approach. The results discussion from line 299 onwards has not been prioritised for space in the manuscript: the authors attempt to concentrate reporting results from six different experiments into a single page, which is not enough exposition for the main text of the article. Many fairly important details are pushed into an Appendix and then covered by asserting the “superior ability of our robust method” in the main text. This is far too simple an analysis for an empirical investigation: I don’t trust any analysis that can be summed up so concisely, let alone six combined, with three being real data sets. Appendix C is reasonably thorough in this respect, although there is a noticeable absence of any rigorous testing of the differences in performance metrics. There are a few small issues with the English in the article: line 8: “among which Ridge…”, line 97: “pervasive” is a strange choice of vocabulary.

Questions

Line 110: “In our setting, instead of a parametric prior on the regression coefficients, we place a nonparametric one on the joint distribution of the response and covariates” What is lost by not having priors associated with parameters? This is not necessarily a good thing to move away from: specific belief distributions for specific interpretable model components is a good thing. What the implications for estimating the posterior uncertainty on the parameter level, for example in the regression case? What kind of posterior convergence do you expect on the parameters themselves? What are the more general implications of putting a prior on the distributions of the data instead of the model parameters? This is suggested to be the case in line 111. Is that not what the likelihood is meant to capture? Is this merely sloppy phrasing in line 111 Does the prior on the distribution of the data come into conflict with the model at some point? What is the difference between altering the prior you have assumed here and altering the statistical model? What is the expected scalability of the methods to larger data sets? None of those used here are really that big, but the use of SGD implies that you have big-N scalability in mind. The wall clock times on your fairly modest desktop system are quite small, so presumably you have some ability to scale to larger data sets if desired. Is the trained DP object interpretable at all? Assuming that the learned DP object implies some sort of partition of the data-generating process, then is that partition telling us something we didn’t know beforehand? How restrictive is the DP prior on the space of measures that you might expect in reality?

Rating

6

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

While an elegant choice for the problem at hand, the use of a Dirichlet Process is likely to be a limiting factor for future directions of research, both in terms of the computational limits and the possible allowable partitions. The assertion of the prior in the way performed here, in contrast to over interpretable parameters, possibly limits the likelihood of this method being widely adopted. The theoretical results are focussed on convergence to the correct expected risk (or transformations thereof), rather than telling us anything about the parameters themselves.

Area Chair vPn42024-08-11

discussion

Dear Reviewer p57R, We appreciate you submitting your review. The authors have provided replies to your comments. Could you kindly let us know if they have adequately addressed your points? Best regards, AC

Reviewer p57R2024-08-12

The authors have engaged with my concerns concerning the article, primarily about the unconventional use of Bayesian reasoning in a new context, and more pragmatic questions of scalability and the presentation of experiments. I appreciate the effort put in. I am happy to keep my score where it is.

Reviewer Vivx7/10 · confidence 4/52024-07-10

Summary

The authors use Dirichelet process and a smooth ambiguity aversion model to approximate the solution of risk minimization problems. They demonstate the consistency of their approach, along with finite sample guarantees. They additionally give a practical means of applying their procedure and apply the procedure to a number of simulated and real problems.

Strengths

Background was very informative and enlightening. Despite being very technical in nature the paper was easy to follow. Based on the references given, the idea seems quite novel. Proofs related to results in section 3 were quite easy to follow (section 4 and appendix proofs not checked)

Weaknesses

There are many small typos. - In Lemma 3.2 the supremum of \gamma^* is not needed as gamma^* does not depend on t. The statement in the proof is the supremum over gamma(t). - For consistency with lemma 3.2, in theorem 3.3 the supremum over gamma(t) should probably be replaced by gamma^* (or gamma(t) should be added back into lemma 3.2). - In line 181 I don't think the (ii) should be there. There are some minor problems. - The introduction of the dependency of phi on n makes a small appearence. This idea is very interesting but I don't think it is emphasised enough. Just a couple of sentences explaining why you'd want to do such a thing would be nice. - In line 221 a form of phi is assumed. I assume this is supposed to be phi_n and not phi. If it is supposed to be phi I'm unsure what form is being assumed, and if it is supposed to be phi_n it doesn't seem like it is assumed as section 4 seems to use phi. And there are some reasonably major (but easily fixable) problems. - On first read it is confusing how exactly the distance between the theoretic risk and the risk computed WRT to p_0 enters into lemma 3.2. Seeing the full statement I understand why the full bound isn't included in the main text but perhaps this discussion can be reworded. - It is unclear if theorem 4.4 is using the same assumptions as lemma 4.3, or just the form of C_T that is given in lemma 4.3. - I think theorem 3.3 needs to assume that M_phi < infinity as it is moved to the otherside of an equation in the proof. I think the proof needs more discussion about the case of gamma^*=infinity as on first read it looks like it moved to another side of the equation in the proof. gamma^*=infinity seems fine as thne the bound is triival. These cases are possible as taking K=1 and phi(t) = (t-1)^(1/3) has M_phi = gamma^*=infinity. Seems like everything can be infinite in lemma 3.2 as then the bound is trivial.

Questions

In theorem 3.4 it is assume that phi_n converges uniformly to the identity function. To me this seems like assuming that curvature of phi_n tends to 0 indictating that as n goes to infinity there is no ambiguity aversion. However, uniform convergence to the identity function is much a much stonger condition than curvature going to 0. For example convergence to 2 times the identity function still has curvature going to 0. Am I wrong in my reasoning behind convergence to the identity function? If not, is it possible to update assumption (3) in theorem 3.4 to a more general curvature to 0 assumption? It seems like the without much modification it can be instead assumed that phi_n converges to some scalar multiple of the idenity function. How do the assumptions on theorem 3.3,3.4 guarantee that theta_n and theta_* exist? As is isn't assume that h is continuous or Theta is compact it seems like the argmin of V and R can be empty. From the proof of theorem 3.3 it seems like if theta_* doesn't exist it can be replaced by any arbitrary theta. Another question. The uniform boundedness of h is strongly relied on to get the generate the results. Is it possible/are there methods available to remove this assumption? Classic problems like least squares regression over an unbounded domain will have h unbounded. This often isn't all that bad as assuming compactness of the domain is fine, but the bounds given here depend strongly on the size of the compact set used to restrict the problem to.

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

4

Limitations

-

Area Chair vPn42024-08-11

discussion

Dear Reviewer Vivx, Thank you for completing your review. The authors have responded to your feedback. Could you please let us know if their responses sufficiently address your concerns? Best regards, AC

Reviewer Vivx2024-08-11

Yes, the response sufficiently addresses my concerns.

Authorsrebuttal2024-08-11

Dear Reviewer, We thank you again for the time spent reading and commenting our work, and remain at your disposal should any further question arise. Best regards, The Authors

Reviewer 7wnT7/10 · confidence 3/52024-07-10

Summary

This paper introduces a new robust risk criterion that integrates concepts from Bayesian nonparametric theory, specifically the Dirichlet process, and a recent decision-theoretic model of smooth ambiguity-averse preferences. They show the relationships between their criterion and traditional regularization methods for empirical risk minimization, including Ridge and LASSO regressions. They also provide theoretical gurantee for the robust optimization and propose tractable approximations of the criterion.

Strengths

This is a solid paper that makes a good and exciting theoretical contribution. The idea of incorporating Dirichlet prior into distributionally robust is novel, and well-suited for the problem. The empirical section is adequate.

Weaknesses

*

Questions

* I have a question about the coefficient in the prior. What will happen if we adopt the empirical Bayesian idea into your algorithm, i.e. estimate the coefficient in the prior through data? Will this cause degradation of the model?

Rating

7

Confidence

3

Soundness

3

Presentation

4

Contribution

3

Limitations

The authors adequately addressed the limitations of their work.

Reviewer uUEe5/10 · confidence 1/52024-07-12

Summary

This paper proposes a novel robust optimization criterion for training machine learning and statistical models by combining Bayesian nonparametric theory and smooth ambiguity-averse preferences, addressing distributional uncertainty to improve out-of-sample performance. The authors demonstrate theoretical guarantees for their method and show its practical implementation and effectiveness through tasks using simulated and real datasets.

Strengths

The theoretical aspects of this paper are not my research area, so I am not equipped to assess the strengths and weaknesses of the theoretical contributions. However, the proposed method and its practical implications appear to be well-founded and promising.

Weaknesses

The theoretical aspects of this paper are not my research area, so I am not equipped to assess the strengths and weaknesses of the theoretical contributions. However, the proposed method and its practical implications appear to be well-founded and promising.

Questions

NO.

Rating

5

Confidence

1

Soundness

3

Presentation

3

Contribution

3

Limitations

Discussed in Section 6.

Reviewer DqoX5/10 · confidence 2/52024-07-23

Summary

The paper proposes a Bayesian method for optimizing stochastic objectives under a given finite sample. More precisely, The paper assumes the Dirichlet Process for data generation and then it proposes to optimize the mean of the stochastic objective over posterior distributions given the sample. The paper proves asymptotic properties of the proposed objective. The experimental results in the appendix show that the proposed method works better than the standard L2 regularization or no regularization on linear regression tasks over simulation and real data sets.

Strengths

The problem is quite relevant. The paper proves basic asymptotic desirable properties. The experimental results show the superiority of the proposed method to the standard L2 regularization-based methods.

Weaknesses

One of my concerns is the proposed method's high computational cost. However, I don’t think it is critical. When the sample size is large enough, the standard ERM would work well. The proposed method, though it shows asymptotic convergence properties, would work better when the sample size is small. On the other hand, it would need more technical improvement when the sample space is high-dimensional.

Questions

Is it difficult to obtain finite-sample error bound like PAC or statistical learning theory? (or any related work?)

Rating

5

Confidence

2

Soundness

3

Presentation

2

Contribution

2

Limitations

N/A

Area Chair vPn42024-08-11

discussion

Dear Reviewer DqoX, Thank you very much for submitting your review report. The author(s) have posted responses to your review. Could you kindly provide comments on whether your concerns have been adequately addressed? Best regards, AC

Area Chair vPn42024-08-11

some comments

Dear Author(s), I would like to invite you to clarify some points which I am uncertain about. Thank you very much! 1. It appears that Propositions 2.1 and 3.1 are well-established, though finding a convenient source for citation might be challenging. I wonder if it would be appropriate to mention this in the revised paper. 2. In comparison with the definition in (1), I wonder if it might be more straightforward to directly link V-hat (as defined in (3)) to ϕ(Rp∗(θ))—the quantity on the right-hand side of the equation in Proposition 3.1—for the analysis. Is there any particular challenge or specific reason for using (1)?

Authorsrebuttal2024-08-11

Dear Area Chair, Thank you for your insightful comments and careful reading of our submission. We’d like to address your questions as follows: 1. We agree with your observation regarding the need for a discussion on the original sources for results similar to Propositions 2.1 and 3.1, and we will incorporate this in the next version of the paper. Specifically, for Proposition 3.1, we will move the citation of [1], currently in the appendix, to the main body as a standard reference on DP weak convergence to the true generating distribution. Regarding Proposition 2.1, we acknowledge that pinpointing the exact origin of the classical interpretation of Ridge regression as Bayesian linear regression with a *parametric* standard normal prior on the coefficients is challenging. However, as noted after Proposition 2.1, our result offers a novel Bayesian interpretation: by taking an optimization-centric view and placing a *nonparametric* prior directly on the data-generating distributions (instead of on the coefficients, which we optimize), we also arrive at Ridge regularization. We believe this result highlights the fundamental connections between optimization, decision theory, and Bayesian inference, and we will ensure this is clearly explained in the revised paper. 2. We also agree that our paper currently lacks an explicit discussion of our reasoning for relating $\hat V$ to $V_{\boldsymbol\xi^n}$ instead of directly to $\phi(\mathcal R_{p_\star}(\cdot))$. We will address this in the next version of the paper with the following clarifications. As you noted, one's ultimate goal may be to ensure the convergence of $\hat V$ to $\phi(\mathcal R_{p_\star}(\cdot))$, which involves three layers of approximation: from the finite sample size $n$, from the random measure truncation threshold $T$, and from the number of MC samples $N$. In Section 3, we address the first layer by studying the convergence of $V_{\boldsymbol\xi^n}$ to $\phi(\mathcal R_{p_\star}(\cdot))$. In Section 4, we focus on the latter two layers, *given a fixed sample size approximation determined* by $n$, by studying the convergence of $\hat V$ to $V_{\boldsymbol\xi^n}$. Thus, within this logical chain, $V_{\boldsymbol\xi^n}$ serves as a bridging quantity between $\hat V$ and $\phi(\mathcal R_{p_\star}(\cdot))$. The results from Sections 3 and 4 can then be combined to directly establish the convergence of $\hat V$ to $\phi(\mathcal R_{p_\star}(\cdot))$, which involves choosing $T$ and $N$ as functions of $n$ to ensure that the right-hand side of the first equation in the statement of Lemma 4.3 converges to zero. We will emphasize this crucial point in our next revision. Thank you again for your valuable feedback, The Authors *[1] S. Ghosal and A. Van der Vaart. Fundamentals of nonparametric Bayesian inference, volume 44. 378 Cambridge University Press, 2017.*

Reviewer 7wnT2024-08-11

Thank you for your rebuttal. I remain positive about the paper.

Authorsrebuttal2024-08-11

Dear Reviewer, We thank you again for the time spent reading and commenting our work, and remain at your disposal should any further question arise. Best regards, The Authors

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC