Detecting hidden confounding in observational data using multiple environments

A common assumption in causal inference from observational data is that there is no hidden confounding. Yet it is, in general, impossible to verify this assumption from a single dataset. Under the assumption of independent causal mechanisms underlying the data-generating process, we demonstrate a way to detect unobserved confounders when having multiple observational datasets coming from different environments. We present a theory for testable conditional independencies that are only absent when there is hidden confounding and examine cases where we violate its assumptions: degenerate&dependent mechanisms, and faithfulness violations. Additionally, we propose a procedure to test these independencies and study its empirical finite-sample behavior using simulation studies and semi-synthetic data based on a real-world dataset. In most cases, the proposed procedure correctly predicts the presence of hidden confounding, particularly when the confounding bias is large.

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer g1qg5/10 · confidence 3/52023-07-06

Summary

The authors develop a test for hidden confounding using only observational data from multiple environments with the same DAG under the independent causal mechanisms assumption.

Strengths

The paper is well-written and covers the related work well. Their setting, assumptions, and contributions are clear and they meaningfully discuss the profound limitations of their work with potential solutions. Under new (but strong) assumptions, they propose a test for no unmeasured confounding and provide an algorithm using only observational data. The proofs in the construction of the test and the test itself is novel and can be counted as original contributions. They support their claims through adequate (semi)-synthetic experiments.

Weaknesses

1. My main concerns about the paper is about its practical utility. - It is not often a practitioner has access to multiple observational datasets - The observational datasets may be highly heterogeneous, e.g., the type of data recorded, the way they are recorded, etc., making the preprocessing task for those datasets to be used in a way that is proposed in this manuscript very hard, if not impossible. - Analyzing observational studies is hard for a lot of reasons. As we increase the number of studies, the chances of introducing bias through other and often overlooked ways increase such as non-adherence to treatment assignments, or selection bias introduced while defining the treatment assignments and where the follow-up begins etc. Those can all be reasons behind rejection beyond confounding. Please correct me if I am wrong in that regard. 2. I find it hard to believe that different observational datasets will have the same DAG. - Different environments (e.g. hospitals) may have very different mechanisms (e.g. institutional practices) that result in different DAGs. 3. In summary, although the paper formalizes a framework where we can test the no unmeasured confounding assumption, I think the approach would only work in very controlled settings.

Questions

I do not have any question beyond some points in the weaknesses part.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The authors acknowledge some important limitations page 6.

Reviewer 7Ey56/10 · confidence 3/52023-07-07

Summary

The authors focus on a setting where we observe treatments, outcomes, and a set of confounders across multiple environments. While it's generally not possible to identify unobserved confounders in a single data set, when we have multiple data sets and assume the mechanisms operate in the same way across those environments (even if the parameters may differ), we can test for the presence of unobserved confounders. The authors propose an independence test between a treatment in one environment and an outcome in another, and show that, if they are not independent (conditioned on the observed confounders and the other treatment), then there must be at least one unobserved confounder that's producing that dependence.

Strengths

The presentation of this paper is excellent, and I found the explanation throughout very clear. The idea comes across as fairly simple, which speaks to the clarity of the explanation. The graphics in particular (Figures 1 and 2) go a long way towards making the approach understandable I appreciate the attention paid to assumptions, including the discussion in 4.1 about assumption relaxation.

Weaknesses

The main weakness I see of this paper is the lack of motivation for the setting the authors consider. This work is focused on a situation where we have observational data from multiple environments where we have collected the same variables from all environments and, while parameters are allowed to differ between them, the actual mechanisms at play need to have the same structure. The authors bring up the example in the intro of different hospitals administering the same treatment but serving different populations, but otherwise, I feel like this work would be served by a discussion of how realistic this sort of setting is, or how common it is in the real world. Similarly, while I do like Section 4.1 and the authors' attention to the assumptions, the lack of clear real-world ties extends here. What would it look like in real data if we had pair-wise dependent mechanisms? Is this something we would expect to see frequently? This lack of discussion makes it hard to assess how important these assumptions actually are, and how useful the relaxations are in turn. Considering that Section 4.2 is the actual core of the approach, I wish more attention were paid to it. I understand that space is limited, but including the actual algorithm you're proposing seems worth moving something else to the appendix.

Questions

While being able to detect whether or not there is hidden confounding is useful, it would be more useful to then make some statement about the strength of hidden confounding, or how much we might expect it to affect our estimates. I suspect in many realistic situations, there is going to be /some/ hidden confounding. If care is taken when choosing which potential confounders to measure, though, the effect of this hidden confounding is hopefully small. Do you think your approach could be used or modified to answer questions about the strength of hidden confounding?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The authors' approach is confined to a fairly specific setting, where the same variables, operating in the same manner (with different parameters), are measured across different environments. The approach also only outputs a binary yes or no for whether hidden confounding exists. If the test comes back positive for hidden confounding, it's unclear if this means causal analysis is impossible, or what the next steps should be. In terms of societal impact, I don't see any real potential for negative impact. This work can be used to test for one potential hurdle for accurate causal inference (hidden confounders), but there are many other assumptions that need to be met as well, and this paper does a good job at making those clear.

Reviewer ZLeH4/10 · confidence 4/52023-07-07

Summary

The authors proposed a detection method for hidden confounders, with only observational data in multiple environments. They, in particular, designed a testable independence condition for it, under the assumptions of (i) Faithfulness & Causal Markov Property, (ii)Shared Causal Graph, (iii) Independent Causal Mechanism Principle, and (iv) Non-degenerate Probabilistic Independent Causal Mechanisms. They also performed synthetic and semi-synthetic experiments to verify the effectiveness of their proposed method.

Strengths

This paper is well written with clear motivation and a description of their method. What the authors focused on is indeed an interesting yet challenging research topic in causal inference. And the idea of using conditional independence conditions, especially between observations, is interesting to me.

Weaknesses

The problem setting has some confusing notations. The experiments cannot verify better performances than other methods for detecting latent confounders.

Questions

Regarding the problem setting, there are some confusing notations. For example, since the $E^{(k)}$consists of four causal mechanisms shown in Fig.1(a), how could it serve as an indicator variable for what environment observations belong to? Btw, why $E^{(k)}$ is observed while the $(\theta_T^{(k)} , \theta_Y^{(k)}, \theta_X^{(k)}, \theta_U^{(k)} )$ are all unobserved, whereas $E = (\theta_T^{(k)} , \theta_Y^{(k)}, \theta_X^{(k)}, \theta_U^{(k)} )$. In line 116, the authors said, “We do not assume to know anything about the individual parameters $(\theta_T, \theta_Y, \theta_X, \theta_U )$, however in their method, they sampled the mechanisms i.i.d. for $(\theta_T, \theta_Y, \theta_X, \theta_U )$. In line 113, ${T, Y, X, Y}$ might be ${T, Y, X, U}$. In Figure 1, $\theta_T, \theta_Y, \theta_X, \theta_U$ are parents of the ${T, Y, X, U}$.respectively. So I suggest removing the subscript $\theta_V$ in Eq.2 since it’s easily misleading for the readers that $\theta_V$ is the parameter of conditional distributions. Regarding the theory, Theorems 1 and 2 seem to derive the independence condition from a fixed environment $k$ other than using the advantages from multiple environments, as shown in Eq.(3). In contrast, the authors used multiple environments for the independent testing. Since the goal is to detect latent confounders, I wonder if some other baselines for a single environment could obtain better results. We might see multiple environments as one environment for this verification. In line 289, the distribution notation for the $\theta_V^{(k)}$ should have reflected the heterogeneity.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

The needed number of environments is so high.

Reviewer 1FZJ7/10 · confidence 4/52023-07-08

Summary

This paper presents a method for detecting hidden confounding in observational data. Without further assumptions, detecting hidden confounding is known to be impossible. This paper takes advantage of a common scenario where one has access to observational data collected from multiple environments. Given particular assumptions that some of the underlying mechanisms and/or exogenous factors are varying across environments, this paper presents a method to detect hidden confounding between a pair of target variables <T,Y>. The paper presents the theory, a discussion of its assumptions and which can be weakened, synthetic experiments and one experiment in a real-world dataset.

Strengths

- detecting hidden confounding under reasonable assumptions is a useful contribution. - the paper is well written and easy to understand. I appreciate the additional detail on assumptions in sec 4.1. - experiments demonstrate implications in synthetic settings and 1 real-world dataset

Weaknesses

- I don't think this approach will handle detection of hidden confounding under selection bias (i.e., where a data collection mechanism means not all data samples are visible in an environment). This is a common challenge in practice, and it is worth mentioning this as a limitation and/or noting it in the initial problem statement. - the paper addresses a scenario where causal mechanisms are parameterized by environment variables $\theta_{T,X,U,V,..}$. If the causal mechanisms themselves are varying, I'm not sure I understand what the motivating causal inference scenario is. That is, if the causal inference question is to identify $P(Y|T)$, how is that a well-defined question if the causal mechanism defining $Y|Pa(Y)$ is changing across environments? - the paper would be improved with additional real world datasets.

Questions

- The current causal graph representation of the problem seems like it cannot represent situations where selection bias is occurring. I.e., where we only observe X,Y,T when certain conditions hold and essentially throw away the data if those conditions do not hold. Is my understanding correct? - Assuming that the parameters $\theta_{T,Y,X,U}$ vary across environments is a strong assumption. I appreciate the discussion of a violation of this assumption starting in line 237. I'd recommend adding a forward-reference to this discussion weakening assumptions. - There are various positivity assumptions here (e.g., in sampling causal mechanisms and in sampling data within an environment). In practice, such positivity assumptions are violated often. E.g., we may gather many sets of data in country A where some variable may take only a small range of values. Later, we may realize there is hidden confounding when we gather new data sets in country B and the given variable takes on a larger range values and, for example, passes some threshold. I am concerned that a naive user of your method will over rely on your method's analysis of data from country A and believe that there is no hidden confounding that will threaten their inferred causal inference mechanism even in country B. Note that this is distinct from the discussion section 4.1 in lines 237-245 --- even if P(T|Pa(T)) is varying --- even if an RCT is performed in country A --- it would not identify that effects will be different in our hypothetical country B. It would be good to highlight this limitation, e.g., as another paragraph in 4.1

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have largely addressed limitations and broader impacts. I appreciate the note of caution regarding use of this work in high-stakes settings.

Reviewer ZLeH2023-08-16

Additional questions

Thanks for your responses. I still have the following concerns. 1. Regarding "Since this condition, while relating two individuals in the same environment, holds for all k, then in a single environment, such two individuals still have this dependence relation." Why have i.i.d. relations? 2. Regarding baselines. Why aren't there options like FCI or similar methods with fewer assumptions considered here? Additionally, there are notable papers that could potentially enrich the discussion and facilitate comparisons. For instance, the following works might warrant attention: Ghassami, A. E., Kiyavash, N., Huang, B., et al. (2018). 'Multi-domain causal structure learning in linear systems.' Advances in Neural Information Processing Systems, 31. Huang, B., Zhang, K., Gong, M., et al. (2020). 'Causal discovery from multiple data sets with non-identical variable sets.' Proceedings of the AAAI Conference on Artificial Intelligence, 34(06), 10153-10161." If there are any misconceptions about the points I've raised above, then I welcome input from the authors.

Authorsrebuttal2023-08-17

Reply to additional questions

Regarding your first question: we interpret it as you asking why we have i.i.d. data within a given environment. The answer is that this follows from our assumptions as, for a given environment $k$, $P(T,Y,X,U \mid \Theta_T^{(k)}, \Theta_Y^{(k)}, \Theta_X^{(k)}, \Theta_U^{(k)})$ is a fixed distribution that we sample independently from. Thus, for a given $k$, the data is i.i.d.. Notice however that we can have dependency relationships between observations in the same environment $k$ when not conditioning on the parameters $\Theta_T^{(k)}, \Theta_Y^{(k)}, \Theta_X^{(k)}, \Theta_U^{(k)}$; this is what we see in Theorem 1 and 2. For the second concern, we thank you for the additional suggestions for baselines. Let's start by addressing why we haven't compared our approach to a standard causal discovery algorithm like FCI. Our primary objective is not to learn a complete causal DAG. Instead, we aim to detect hidden confounding while having some prior knowledge of the underlying structure. Additionally, although FCI is a natural consideration, it does not work in our context. With the ground-truth graph in Figure 2 from our paper, there are no observed conditional independencies. Thus, FCI outputs an uninformative PAG that gives no information on whether we can exclude the presence of a hidden confounder between $T$ and $Y$. As we discuss in section C in the Appendix of our paper, there are graphs similar to the one in Figure 2 where ordinary constraint-based algorithms might be able to detect hidden confounding, but these do not cover all possible cases of our problem setting. Further, regarding the papers you mentioned: both are relevant to our work and deserve inclusion in our related works section, which we will address. For instance, Ghassami et al. (2018) also leverage the independent causal mechanism principle for structure learning. However, neither of these papers are considering the same problem as us. Ghassami et al. (2018) assume, in contrast to us, that all variables are observed (i.e. causal sufficiency) which means that their method can not be used for detecting hidden confounding. Meanwhile, Huang et al. (2020) assume constant linear relationships across environments: $b_{ij}$ appears fixed in (1) in their paper. In addition to assuming linearity unlike us, this constant relationship assumption represents a stringent constraint not imposed within our framework.

Reviewer g1qg2023-08-19

Response to Rebuttal

I thank the authors for their response. I have read the other reviews and responses as well, and decided to keep my original score.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC