Summary
The paper studies a linear Structural Causal Model (SCM) for prediction under distribution shifts due to an an additive term to covariates that changes at test time, and an unobserved confounder between the covariates and the label. The key difference between the proposed analysis and those presented in other works (e.g. those in anchor regression, invariant causal prediction etc.) is that the shift is bounded in its strength. Another difference is that the optimal robust predictor w.r.t the entire uncertainty set is not identifiable. Hence, instead of the optimal robust error, the paper studies the optimal error that can be achieved from observable data. This corresponds to the robust error over an uncertainty set that is generated by all SCM parameters that could have produced the training data, which is in general a larger set than the uncertainty set we are interested in (i.e. of bounded additive shifts to the covariates).
Once the problems is set up, a parametric form for the set of parameters that can generate the training data is derived, along with the corresponding set of possible robust predictors. Then under a mild assumption on boundedness of the ground truth regression parameters, the following are derived: 1) a closed form is derived for the robust loss in the case where test distributions only induce shifts in directions that have been observed at training. The loss grows linearly with the strength of the shift; 2) a lower bound on the achievable risk from observed data, which is tight (and thus can be learned from observed data) for large enough shifts.
Further analysis and simulations with Gaussian data are done to compare the risk of the derived estimator with that of ERM and anchor regression (which does not exploit the boundedness of shifts to achieve better prediction). The results verify that the risk obtained by the proposed empirical estimator are close the lower bound given in the theorem, and outperform the two baselines as shifts grow large.
Strengths
* Overall I enjoyed reading the paper, it is clear and written with care. Generally, I also like the direction of formalizing bounded shifts and unidentifiable settings in detail. Finally, the analysis is well performed and easy to follow.
* The work is original in formally analyzing bounded distribution shifts, where even in population the optimal robust predictor might still have non-zero projection on directions that shift at test time. \
Small note on this: I think that for finite sample guarantees boundedness assumptions on the strength of the shift must be made, and they are made in works that give sample complexity results. I believe the reason is that with unbounded shifts, any small weight on a shifting feature can be magnified unboundedly to yield a large robust error. With finite samples, it is usually impossible to guarantee strictly zero weights on the shifting directions. Therefore it might be worthwhile clarifying that considering the population loss is an important component for the analysis.
* Beyond the points mentioned above, the bounds on the robust risk, closed form solutions, and the formalism used in the work, can be useful for future theoretical analyses. In terms of practical significance, the approach derived from these results to explicitly account for bounded shifts might be useful in the future, if it is generalized beyond linear models.
Weaknesses
* As mentioned above and the authors mention in describing the limitations of their work, it is restricted to linear models. Intuitively, a considerable challenge in learning robust models is to identify the ``directions” that shift between domain. Under linear models it is rather straightforward, and the method proposed in the paper depends on linearity of the model in order to find these directions. This is unlike some other methods in the domain generalization literature, where certain formal results are provided on linear models but the methods can easily be tested in the non-linear case.
* Even within the realm of linear models, I am not throughly convinced that the method is useful in real world problems. To make this more convincing, it might have been nice to run simulations over non-synthetic data and some more variants of methods. E.g. using several methods that learn invariant models when hyper-parameter tuning is performed with the objective proposed in this work may be of interest. That is to see whether boundedness of the shift should be taken into account during training, or is it enough to simply use it for model selection. Yet the most significant drawback is still of course, the limitation to linear models.
Some other/smaller comments:
* In line 60 prior work is cited to claim "that even minor violations of the identifiability assumptions can cause invariance based methods to perform equally or worse than ERM". However, I am not sure that the failures portrayed in these works are strictly due to identifiability violations. At least not the ones alluded to in this paper, which is the lack of heterogeneity in the training environments. In Kamath et al. 21, the predictor is identifiable but the problem is specifically in the IRMv1 objective, see also [1] who point out an objective that solves this issue. In Rosenfeld et al. 20, only one failure is due to not having enough environments (and arguably, this is since they consider the uncertainty set that includes shifts in all directions. I will also touch on this in the next minor comment), and the second one is due to non-linearity.
* Regarding analysis under lack of heterogeneity or unidentifiable optimal robust predictor: I think that the part of the analysis that touches upon the insufficient heterogeneity of the environments in order to identify the robustness set/robust predictor can be slightly reframed. If I am not mistaken, even in works that give results about identifiable robustness sets (e.g. when the number of environments is linear in the number of shifting features, or stronger ones like [2]), it is most likely possible to draw guarantees about robust risks, but only with respect to a smaller uncertainty set which is restricted to dimensions that shifted between the training environments. That is since the methods still enforce constraints which restrict weights in certain directions. While it is true that most of these prior works did not consider unidentifiable uncertainty sets, it might be related to choice of presentation, and not strictly because the methods are technically limited in that sense. Further, in lines 168-169 it is claimed that “prior work only considers scenarios where the robustness set and hence also the robust prediction model are still identifiable”. I am not sure this is entirely true. There are different formalisms that were considered for spurious correlations, such as those in [3]. In their setting, when the association is not what they call “purely spurious”, then the robust predictor cannot necessarily be recovered. Also [4] study a similar setting where due to similar reasons no guarantee can be given on identifying the robust predictor, but only a lower bound on the error is given.
* It might be worth mentioning in lines 112-113 that in case $H$ in figure 1 is a selection variable (I assume it can be since both edges are bi-directional) then the shift reduces to covariate shift, if I am not mistaken.
* Personally, I found the notation in the introduction which uses $\mathrm{shift}\in{\mathrm{shift set}}$ etc. somewhat redundant and confusing. The more formal notation in section 2 onwards was much easier to understand, so in my opinion it might be worth considering to go straight ahead into that notation.
[1] Wald Y, Feder A, Greenfeld D, Shalit U. On calibration and out-of-domain generalization. Advances in neural information processing systems. 2021 Dec 6;34:2215-27.
[2] Chen Y, Rosenfeld E, Sellke M, Ma T, Risteski A. Iterative feature matching: Toward provable domain generalization with logarithmic environments. Advances in Neural Information Processing Systems. 2022 Dec 6;35:1725-36.
[3] Veitch V, D'Amour A, Yadlowsky S, Eisenstein J. Counterfactual invariance to spurious correlations in text classification. Advances in neural information processing systems. 2021 Dec 6;34:16196-208.
[4] Puli AM, Zhang LH, Oermann EK, Ranganath R. Out-of-distribution Generalization in the Presence of Nuisance-Induced Spurious Correlations. In International Conference on Learning Representations.
Questions
* It could’ve been nice to derive upper bounds on the robust loss in addition to the lower bounds proved in the theorem. That is, assuming these bounds can be calculated from observable data, I’d imagine that most practitioners would be more interested in an upper bound as it gives a strong guarantee on the risk worst possible risk. Do you think this is a reasonable goal, and do you perhaps have any insights on this?
* Could you perhaps discuss the differences between the empirical optimization problem you derived in the paper and those derived in prior work? For instance, looking at eq. 19 in appendix D, it seems like there is a term that penalizes $\|\| \hat{S}^\top(\hat{\beta}^{\mathcal{S}} - \hat{\beta}) \|\|$ which I read as the weights in the “invariant directions” should be the same across the environments. This seems quite similar to ICP, IRM etc. Then the other term limits the projection on the shifting directions, but it might be useful for readers to have a short discussion on the conceptual similarities between this and penalized versions of other methods. Also, there might be a typo in that equation, $\mathrm{arg}\min_{\beta\in{\mathbb{R}^d}}$ should be $\mathrm{arg}\min_{\hat{\beta}\in{\mathbb{R}^d}}$ instead.
Limitations
The authors have adequately discussed limitations.