Achievable distributional robustness when the robust risk is only partially identified

In safety-critical applications, machine learning models should generalize well under worst-case distribution shifts, that is, have a small robust risk. Invariance-based algorithms can provably take advantage of structural assumptions on the shifts when the training distributions are heterogeneous enough to identify the robust risk. However, in practice, such identifiability conditions are rarely satisfied -- a scenario so far underexplored in the theoretical literature. In this paper, we aim to fill the gap and propose to study the more general setting when the robust risk is only partially identifiable. In particular, we introduce the worst-case robust risk as a new measure of robustness that is always well-defined regardless of identifiability. Its minimum corresponds to an algorithm-independent (population) minimax quantity that measures the best achievable robustness under partial identifiability. While these concepts can be defined more broadly, in this paper we introduce and derive them explicitly for a linear model for concreteness of the presentation. First, we show that existing robustness methods are provably suboptimal in the partially identifiable case. We then evaluate these methods and the minimizer of the (empirical) worst-case robust risk on real-world gene expression data and find a similar trend: the test error of existing robustness methods grows increasingly suboptimal as the fraction of data from unseen environments increases, whereas accounting for partial identifiability allows for better generalization.

Paper

References (63)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer 2y9v7/10 · confidence 1/52024-07-10

Summary

This paper proposed a general framework of partially identifiable robustness to evaluate robustess in scenarios where the training distributions are not heterogeneous enough to identify the robust risk. They define 'the identifiable robust risk' and its correspondig minimax quantity. They show previous approaches achieve suboptimal robustness in this scenario. Finally, they propose the empirical minimizer of the identifiable robust risk and show that it outperforms existing methods in finite-sample experiments.

Strengths

The paper is crealy written. The idea of establishing partially identifiability to fill the gap between an "all-or-nothing" view on robustness in the SCM framework is interesting.

Weaknesses

I am not an expert in this field so I cannot point out many weaknesses. Maybe one thing is the lack of motivation from real-world application, such as in what applications or when the proposed framework is usefull in practice?

Questions

1. On line 135-137, the $M_{\text{test}}$ is defined as $M_{\text{test}}=\gamma\Pi_{M}$. What's the reasoning behind this choice?

Rating

7

Confidence

1

Soundness

3

Presentation

3

Contribution

3

Limitations

yes

Reviewer faJC7/10 · confidence 3/52024-07-13

Summary

This paper investigates the optimal minimax risk of a robust predictor when the robustness set is partially observable, under a structural causal model with hidden confounders. By decomposing the test covariance matrix of latent parameters into a component spanned by the training distributions and the orthogonal component, it is shown that the robust predictor is only identifiable if the test shifts are in the direction of the training shifts. While for unobserved test shifts, the best achievable minimax risk grows linearly with shift scale. The theory is applied to show sub-optimality of OLS and anchor regression with partially observable test shifts.

Strengths

I am not familiar with the field of causal inference, but it appears from the paper that it has made contributions as the first result for distributional robustness or invariant causal prediction when the robustness set is partially identifiable. Particularly, it shows that infinite robustness is impossible under this scenario, and finite robustness methods can show performances reduced to ERM. This echoes empirical evidence and provides a possible theoretical explanation for reported failure of distributional robustness methods under wild environments.

Weaknesses

The structural causal model in Eq 2 and the resulting explicit solution in Eq 6 shows that the model is biased even without distribution shifts. To see this, take $\gamma =0$ in Eq 6 and the predictor does not reduce to $\beta^\star$. The model is unbiased only when the cross covariance between $\eta$ and $\xi$ vanishes, which implies no hidden confounder and only covariate shift. In contrast, the classic approach for invariant causal prediction [1] produces unbiased estimation beyond covariate shift. [1] Arjovsky, M., Bottou, L., Gulrajani, I., & Lopez-Paz, D. (2019). Invariant risk minimization. arXiv preprint arXiv:1907.02893.

Questions

1. Typo: L137, the sub-space. 2. In Fig 1, bidirectional edges don't make sense for causal graphs, and it's not explained either. Moreover, is there exact equivalence between Fig 1 and the SCM in Eq 2?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The author has addressed limitations with regarding to the model, e.g., linear structure and additive noises.

Reviewer iX3g7/10 · confidence 4/52024-07-14

Summary

This paper proposes a new framework for distributionally robustness under the linear causal setting. Specifically, the authors minimize the so-called identifiable robust risk, which corresponds to the maximum of the robust risk for parameters in the observationally equivalent set. Under such partially identifiable robustness framework, they discuss the lower bound of the risk and show that some estimated identifiable robust predictor can achieve that lower bound risk, where the corresponding empirical version can approximate the lower bound risk well. They also validate the better performance with finite sample numerical experiments.

Strengths

- The motivating illustration and results are very clear; - The direction the authors study is interesting in bridging the structure and non-structure of distributional robustness.

Weaknesses

- Some paragraph in the training and testing data part can be reorganized as some formal assumptions, such as the assumption of M_{test} and linear assumption as well as the discussion in Section 3.2 in terms of the definition of S and M. This gives a clear picture on what some key conditions (and possible relaxations) are in this paper. - There are a few missing details (see Questions as follows), which I hope the authors can explain them a bit more.

Questions

- Can the authors highlight a bit more on the (approximated) real-world example when the predictor is partially identifiable? Since it may be a bit unfamiliar to the general DRO audience for these notations and their relevance in practice. - Can the authors elaborate more on how the empirical estimations for the space of the training shift direction $\hat S$ and $\hat R$ are computed, as compared with the current version in Appendix D? From my preliminary understanding, this space determination is important for subsequent estimation. - The potential utility of **active intervention selection** is also interesting, and it seems aligned with some recent relevant work on distribution shifts. For example, in causal explanation [1] and incorporating specific features in improving distributional robustness [2]. I am wondering if the authors can provide more details on how their own intervention is and the connections with these existing literature, which can be potentially incorporated in the main body and appendix. [1] Quintas-Martinez, Victor, et al. "Multiply-Robust Causal Change Attribution." arXiv preprint arXiv:2404.08839 (2024). [2] Liu, Jiashuo, et al. "On the Need for a Modeling Language Describing Distribution Shifts: Illustrations on Tabular Datasets." arXiv preprint arXiv:2307.05284 (2023).

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors discuss the main limitation of this paper, which I think is reasonable compared with existing literature.

Reviewer PdW86/10 · confidence 4/52024-07-29

Summary

The paper studies a linear Structural Causal Model (SCM) for prediction under distribution shifts due to an an additive term to covariates that changes at test time, and an unobserved confounder between the covariates and the label. The key difference between the proposed analysis and those presented in other works (e.g. those in anchor regression, invariant causal prediction etc.) is that the shift is bounded in its strength. Another difference is that the optimal robust predictor w.r.t the entire uncertainty set is not identifiable. Hence, instead of the optimal robust error, the paper studies the optimal error that can be achieved from observable data. This corresponds to the robust error over an uncertainty set that is generated by all SCM parameters that could have produced the training data, which is in general a larger set than the uncertainty set we are interested in (i.e. of bounded additive shifts to the covariates). Once the problems is set up, a parametric form for the set of parameters that can generate the training data is derived, along with the corresponding set of possible robust predictors. Then under a mild assumption on boundedness of the ground truth regression parameters, the following are derived: 1) a closed form is derived for the robust loss in the case where test distributions only induce shifts in directions that have been observed at training. The loss grows linearly with the strength of the shift; 2) a lower bound on the achievable risk from observed data, which is tight (and thus can be learned from observed data) for large enough shifts. Further analysis and simulations with Gaussian data are done to compare the risk of the derived estimator with that of ERM and anchor regression (which does not exploit the boundedness of shifts to achieve better prediction). The results verify that the risk obtained by the proposed empirical estimator are close the lower bound given in the theorem, and outperform the two baselines as shifts grow large.

Strengths

* Overall I enjoyed reading the paper, it is clear and written with care. Generally, I also like the direction of formalizing bounded shifts and unidentifiable settings in detail. Finally, the analysis is well performed and easy to follow. * The work is original in formally analyzing bounded distribution shifts, where even in population the optimal robust predictor might still have non-zero projection on directions that shift at test time. \ Small note on this: I think that for finite sample guarantees boundedness assumptions on the strength of the shift must be made, and they are made in works that give sample complexity results. I believe the reason is that with unbounded shifts, any small weight on a shifting feature can be magnified unboundedly to yield a large robust error. With finite samples, it is usually impossible to guarantee strictly zero weights on the shifting directions. Therefore it might be worthwhile clarifying that considering the population loss is an important component for the analysis. * Beyond the points mentioned above, the bounds on the robust risk, closed form solutions, and the formalism used in the work, can be useful for future theoretical analyses. In terms of practical significance, the approach derived from these results to explicitly account for bounded shifts might be useful in the future, if it is generalized beyond linear models.

Weaknesses

* As mentioned above and the authors mention in describing the limitations of their work, it is restricted to linear models. Intuitively, a considerable challenge in learning robust models is to identify the ``directions” that shift between domain. Under linear models it is rather straightforward, and the method proposed in the paper depends on linearity of the model in order to find these directions. This is unlike some other methods in the domain generalization literature, where certain formal results are provided on linear models but the methods can easily be tested in the non-linear case. * Even within the realm of linear models, I am not throughly convinced that the method is useful in real world problems. To make this more convincing, it might have been nice to run simulations over non-synthetic data and some more variants of methods. E.g. using several methods that learn invariant models when hyper-parameter tuning is performed with the objective proposed in this work may be of interest. That is to see whether boundedness of the shift should be taken into account during training, or is it enough to simply use it for model selection. Yet the most significant drawback is still of course, the limitation to linear models. Some other/smaller comments: * In line 60 prior work is cited to claim "that even minor violations of the identifiability assumptions can cause invariance based methods to perform equally or worse than ERM". However, I am not sure that the failures portrayed in these works are strictly due to identifiability violations. At least not the ones alluded to in this paper, which is the lack of heterogeneity in the training environments. In Kamath et al. 21, the predictor is identifiable but the problem is specifically in the IRMv1 objective, see also [1] who point out an objective that solves this issue. In Rosenfeld et al. 20, only one failure is due to not having enough environments (and arguably, this is since they consider the uncertainty set that includes shifts in all directions. I will also touch on this in the next minor comment), and the second one is due to non-linearity. * Regarding analysis under lack of heterogeneity or unidentifiable optimal robust predictor: I think that the part of the analysis that touches upon the insufficient heterogeneity of the environments in order to identify the robustness set/robust predictor can be slightly reframed. If I am not mistaken, even in works that give results about identifiable robustness sets (e.g. when the number of environments is linear in the number of shifting features, or stronger ones like [2]), it is most likely possible to draw guarantees about robust risks, but only with respect to a smaller uncertainty set which is restricted to dimensions that shifted between the training environments. That is since the methods still enforce constraints which restrict weights in certain directions. While it is true that most of these prior works did not consider unidentifiable uncertainty sets, it might be related to choice of presentation, and not strictly because the methods are technically limited in that sense. Further, in lines 168-169 it is claimed that “prior work only considers scenarios where the robustness set and hence also the robust prediction model are still identifiable”. I am not sure this is entirely true. There are different formalisms that were considered for spurious correlations, such as those in [3]. In their setting, when the association is not what they call “purely spurious”, then the robust predictor cannot necessarily be recovered. Also [4] study a similar setting where due to similar reasons no guarantee can be given on identifying the robust predictor, but only a lower bound on the error is given. * It might be worth mentioning in lines 112-113 that in case $H$ in figure 1 is a selection variable (I assume it can be since both edges are bi-directional) then the shift reduces to covariate shift, if I am not mistaken. * Personally, I found the notation in the introduction which uses $\mathrm{shift}\in{\mathrm{shift set}}$ etc. somewhat redundant and confusing. The more formal notation in section 2 onwards was much easier to understand, so in my opinion it might be worth considering to go straight ahead into that notation. [1] Wald Y, Feder A, Greenfeld D, Shalit U. On calibration and out-of-domain generalization. Advances in neural information processing systems. 2021 Dec 6;34:2215-27. [2] Chen Y, Rosenfeld E, Sellke M, Ma T, Risteski A. Iterative feature matching: Toward provable domain generalization with logarithmic environments. Advances in Neural Information Processing Systems. 2022 Dec 6;35:1725-36. [3] Veitch V, D'Amour A, Yadlowsky S, Eisenstein J. Counterfactual invariance to spurious correlations in text classification. Advances in neural information processing systems. 2021 Dec 6;34:16196-208. [4] Puli AM, Zhang LH, Oermann EK, Ranganath R. Out-of-distribution Generalization in the Presence of Nuisance-Induced Spurious Correlations. In International Conference on Learning Representations.

Questions

* It could’ve been nice to derive upper bounds on the robust loss in addition to the lower bounds proved in the theorem. That is, assuming these bounds can be calculated from observable data, I’d imagine that most practitioners would be more interested in an upper bound as it gives a strong guarantee on the risk worst possible risk. Do you think this is a reasonable goal, and do you perhaps have any insights on this? * Could you perhaps discuss the differences between the empirical optimization problem you derived in the paper and those derived in prior work? For instance, looking at eq. 19 in appendix D, it seems like there is a term that penalizes $\|\| \hat{S}^\top(\hat{\beta}^{\mathcal{S}} - \hat{\beta}) \|\|$ which I read as the weights in the “invariant directions” should be the same across the environments. This seems quite similar to ICP, IRM etc. Then the other term limits the projection on the shifting directions, but it might be useful for readers to have a short discussion on the conceptual similarities between this and penalized versions of other methods. Also, there might be a typo in that equation, $\mathrm{arg}\min_{\beta\in{\mathbb{R}^d}}$ should be $\mathrm{arg}\min_{\hat{\beta}\in{\mathbb{R}^d}}$ instead.

Rating

6

Confidence

4

Soundness

4

Presentation

3

Contribution

3

Limitations

The authors have adequately discussed limitations.

Reviewer faJC2024-08-07

I acknowledge and thank the author for the response. In my review, my concern over the usefulness of the structural causal model is addressed. I recognize that the non-existence of an unbiased causal effect estimator for infinite robustness is expected where bounded distribution shift is considered, which is also part of the contribution. Therefore, I raise my score from 6 (weak accept) to 7 (accept).

Reviewer iX3g2024-08-09

I acknowledge and thank the authors for their detailed response. Therefore, I keep my evaluation towards the paper. Specificially, I am quite interested in the active intervention part authors mention and agree with what the authors say in the rebuttal with respect to its connection with the active learning side. Therefore, I am eager to see their revised version. Btw, for [3] (i.e. https://arxiv.org/pdf/2307.05284), I am pointing out their intervention part (i.e., Section 5 in their latest arXiv version instead of the conference version), which also discusses some potential feature-based interventions and can be of separate utility to the authors.

Reviewer PdW82024-08-10

Thank you for addressing the comments and questions, I also appreciate the additional empirical results. I see the point in most of the replies, perhaps one small comment is that [3, 4] indeed provide some binary negative statements, but [4] also give a form of a guarantee (an estimator that has better-than-random worst case performance, and is optimal in a restricted manner). I agree that it is a different flavor of guarantee and analysis than the one given in this paper. Hence my comment is just meant to suggest a slight adjustment to the framing of the contribution, not to devalue it in any way. I have raised my score following the in-depth response.

Reviewer 2y9v2024-08-11

Thanks for the reply. It addressed my questions.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC