Summary
The authors proposed a detection method for hidden confounders, with only observational data in multiple environments. They, in particular, designed a testable independence condition for it, under the assumptions of (i) Faithfulness & Causal Markov Property, (ii)Shared Causal Graph, (iii) Independent Causal Mechanism Principle, and (iv) Non-degenerate Probabilistic Independent Causal Mechanisms. They also performed synthetic and semi-synthetic experiments to verify the effectiveness of their proposed method.
Strengths
This paper is well written with clear motivation and a description of their method.
What the authors focused on is indeed an interesting yet challenging research topic in causal inference. And the idea of using conditional independence conditions, especially between observations, is interesting to me.
Questions
Regarding the problem setting, there are some confusing notations. For example, since the $E^{(k)}$consists of four causal mechanisms shown in Fig.1(a), how could it serve as an indicator variable for what environment observations belong to? Btw, why $E^{(k)}$ is observed while the $(\theta_T^{(k)} , \theta_Y^{(k)}, \theta_X^{(k)}, \theta_U^{(k)} )$ are all unobserved, whereas $E = (\theta_T^{(k)} , \theta_Y^{(k)}, \theta_X^{(k)}, \theta_U^{(k)} )$.
In line 116, the authors said, “We do not assume to know anything about the individual parameters $(\theta_T, \theta_Y, \theta_X, \theta_U )$, however in their method, they sampled the mechanisms i.i.d. for $(\theta_T, \theta_Y, \theta_X, \theta_U )$.
In line 113, ${T, Y, X, Y}$ might be ${T, Y, X, U}$.
In Figure 1, $\theta_T, \theta_Y, \theta_X, \theta_U$ are parents of the ${T, Y, X, U}$.respectively. So I suggest removing the subscript $\theta_V$ in Eq.2 since it’s easily misleading for the readers that $\theta_V$ is the parameter of conditional distributions.
Regarding the theory, Theorems 1 and 2 seem to derive the independence condition from a fixed environment $k$ other than using the advantages from multiple environments, as shown in Eq.(3). In contrast, the authors used multiple environments for the independent testing.
Since the goal is to detect latent confounders, I wonder if some other baselines for a single environment could obtain better results. We might see multiple environments as one environment for this verification.
In line 289, the distribution notation for the $\theta_V^{(k)}$ should have reflected the heterogeneity.
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.