Nonparametric Identifiability of Causal Representations from Unknown Interventions

We study causal representation learning, the task of inferring latent causal variables and their causal relations from high-dimensional mixtures of the variables. Prior work relies on weak supervision, in the form of counterfactual pre- and post-intervention views or temporal structure; places restrictive assumptions, such as linearity, on the mixing function or latent causal model; or requires partial knowledge of the generative process, such as the causal graph or intervention targets. We instead consider the general setting in which both the causal model and the mixing function are nonparametric. The learning signal takes the form of multiple datasets, or environments, arising from unknown interventions in the underlying causal model. Our goal is to identify both the ground truth latents and their causal graph up to a set of ambiguities which we show to be irresolvable from interventional data. We study the fundamental setting of two causal variables and prove that the observational distribution and one perfect intervention per node suffice for identifiability, subject to a genericity condition. This condition rules out spurious solutions that involve fine-tuning of the intervened and observational distributions, mirroring similar conditions for nonlinear cause-effect inference. For an arbitrary number of variables, we show that at least one pair of distinct perfect interventional domains per node guarantees identifiability. Further, we demonstrate that the strengths of causal influences among the latent variables are preserved by all equivalent solutions, rendering the inferred representation appropriate for drawing causal conclusions from new data. Our study provides the first identifiability results for the general nonparametric setting with unknown interventions, and elucidates what is possible and impossible for causal representation learning without more direct supervision.

Paper

Similar papers

Peer review

Reviewer ZLAN6/10 · confidence 4/52023-06-29

Summary

The paper discusses the task of identifying causal variables from high-dimensional observations under non-parametric mixing functions and causal mechanisms. This is done under the assumption of single-node, perfect interventions being available for all causal variables, as well as distinct paired perfect interventions in the case of having more more than two causal variables. The paper proves that causal variables are identifiable under this setup, when taking additional assumptions on the interventions being sufficiently different from the observational distribution. Thereby, weaker assumptions are possible for 2 variables than for more variables. Finally, the paper sketches possible implementations of learning algorithms for this setting.

Strengths

The paper is overall well written, even if it is aimed at researchers in identifiability and/or causal representation learning (CRL) specifically. A consistent notation is used throughout the paper, and all assumptions are clearly stated before the theorems. It is appreciated that proof sketches have been included in the main paper to support the claimed theorems and make the main paper a bit more standalone. The paper discusses all necessary related work and puts itself into context of the current field of research. The main contribution of the paper is its theoretical result. It extends the domain of identifiable causal representation by considering yet another setup, where environment pairs with single-node, perfect interventions are given. The benefit of this setup is that it does not require counterfactual observations, while supporting a large function class despite taking needed assumptions on the interventions. The proofs for supporting the claimed theorems are given in the appendix, following common proof strategies in CRL. The proofs appear sound and intuitive, although a very careful check of the proofs was not possible during the review period. Overall, it is a good contribution to the theoretical identifiability in CRL.

Weaknesses

While the derived theory in the paper puts weaker constraints on the mechanisms of the causal variables and mixing function, its assumption of having access to single-node, perfect interventions on all causal variables is restrictive. Being able to perform an intervention on a variable is already commonly considered expensive or often not easily feasible, especially if it is a perfect intervention and single-node. However, doing this twice and even different between the two setups is challenging. Further, obtaining such a dataset requires non-trivial prior knowledge of the causal system, since it necessitates the ability to perform such single-node, perfect interventions on causal variables that are yet to be identified. The paper misses to give real-world examples to motivate the setup and its assumptions, which puts it in a more limited spot. Besides the theoretical results, it is also important to validate the setup and the practicality of the theory in empirical studies. The paper only sketches some potential ideas, where all unknown parts are learned. However, optimizing the latent encoder, the causal graph, and the intervention targets all at the same time is not trivial as shown in previous works. Further, the appendix shows some limited results on a generative model, where one would need to iterate over all possible causal graphs and intervention targets. Still, this is not practical for systems larger than very few causal variables or high-dimensional observations. The paper states that the intervention targets are not known. However, under the identifiability up to permutation, the intervention targets in this setup appear to be known. Specifically, assumption (A2') states that there exist $n$ environment pairs, where each pair intervenes on a different causal variable. Thus, the intervention targets for these pairs, as stated in the assumption, are known as $\pi(i)$. Since the variables cannot be identified up to permutation $\pi$ anyway, permuting the causal variables and thus the targets are still considered to be the same targets in the same identifiability class, e.g. as in the works cited for known intervention targets [69, 70]. Thus, the claim of unknown intervention targets appears not valid given the assumptions, or the assumptions should be clarified to e.g. have at least $n$/$n+1$ environments. ### Typos: - Table 1: 'Causal Representation Learning'

Questions

### Review summary The theoretical results of the paper provide a new setting under which causal variables are identifiable in CRL. However, the paper is limited by its strong reliance on single-node, perfect interventions and very limited empirical study. I consider the theoretical results outweighing the drawbacks a bit, although the paper would strongly benefit from empirical validation of the setup. Thus, my recommendation is 'Weak Accept'. ### Questions - What is a real-world scenario in which the setup of pairs of single-node, perfect interventions is practical and common? - Do you require the knowledge of the intervention targets up to permutation, or do you allow for more environments/environment pairs as long as each variable has been intervened upon once?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Limitations have been discussed in different parts of the paper.

Reviewer E44d6/10 · confidence 2/52023-07-03

Summary

The paper studies the problem of inferring causal relationships between $n$ latent variables through observations under a mixing function. Given data $X$ from multiple environments (each of which corresponds to an unknown perfect atomic intervention), where $X$ is the observation of the latents under a fixed mixing function $f$, the goal is to recover $f^{-1}$ and the causal graph $G$ on the latent variables (up to $\sim_{CRL}$ equivalence).

Strengths

The problem is well-motivated and interesting. The authors did a good job explaining how this paper differs from prior work while providing a pretty good literature review. Some experiments are also given in Appendix D.

Weaknesses

While I am not an expert in the area and did not check all the proofs in detail, I do not see any glaring weaknesses. The theorem statements and proof sketches seem believable, especially since there is a lot of assumptions that were made to "make things go through". My biggest gripe is that there is a lack of discussion about the assumptions (see Questions section).

Questions

Line 175: Maybe write what CRL stands for somewhere (possibly in the footnotes)? Assumptions: There are a lot of assumptions (which is okay, if they are well-justified and discussed). Can you explain or discuss why each of them is necessary or reasonable to have (without trivializing the problem)? I understand that you believe "pairs of environments" is not necessary in general, but what about the other assumptions? What happens if all but one is satisfied? What goes wrong? I am happy to further increase my score if this is sufficiently addressed and if the other reviewers did not raise any damning issues that I missed. Table 1 caption: Typo: "Reresentation Learning"

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Nil.

Reviewer U76e7/10 · confidence 3/52023-07-05

Summary

This paper proposed to identify the latent causal representations and their underlying causal structure, which is a very challenging and interesting problem. The \sim_{CRL} is introduced to describe the equivalent class up to elementwise operations and permutation, which is sufficiently meaningful for practical use. The CRL-identifiability theory is given under the data from paired interventional data and other assumptions, such as the pre-given number of nodes and others. In appendix, the authors presented a simple version of learning method and validates it on a synthetic dataset.

Strengths

In general, I found this paper to be highly enjoyable and insightful. It successfully addresses a challenging and captivating problem of extracting causal representations and their relationships. Given the increasing prevalence of unstructured data, such an endeavor holds significant importance. The authors have provided a comprehensive overview and engaging discussions that effectively highlight the unique contributions of their work in relation to existing literature. Moreover, the use of paired interventional data, which is more readily obtainable in practical scenarios, adds to the paper's practical relevance. Besides, the organization and writing of this paper are commendable.

Weaknesses

1. I recommend that the authors provide practical demonstrations of the proposed method in real-world scenarios. While acquiring paired interventional data can be challenging in real-world settings, the authors could consider utilizing datasets generated from virtual environments, such as the causal world (https://sites.google.com/view/causal-world/home), to showcase the utility of their approach. 2. In practical applications, determining the number of latent nodes n, is often difficult. Consequently, verifying whether the number of paired intervention data includes all latent variables becomes challenging. This limitation may restrict the scope of application for the proposed theory and learning methods.

Questions

1. Intuitively, it appears that with paired interventional data, we can identify the latent representation with a unique permutation. For example, if we intervene to place a ball at position A on the table in e_i and at position B on the table in e'_i, we can infer that the position is the intervened variable through a simple comparison. Could you please provide a more in-depth explanation of the challenges encountered in theoretical analysis? 2. The requirement of the genericity condition is specified in Theorem 4.1, whereas it is not explicitly mentioned in Theorem 4.2, which is presented as a more general version of Theorem 4.1. Additionally, the assumption A_2' does not degenerate to A_2 when n=2; instead, it is stronger than extending A_2 to a general n. I would appreciate a more detailed clarification regarding this matter.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

Yes, the authors adequately addressed the limitations.

Reviewer FsUF6/10 · confidence 3/52023-07-10

Summary

This paper gives an identifiability result in a setting that is relevant to causal representation learning, where we wish to infer latent causal variables and their causal graph from high-dimensional observations. They work in a setting that is more general than prior work that relies on, for example, weak supervision, temporal structure, or known intervention targets. Their setting assumes that both the causal model and the mixing function are nonparametric, and the targets of the interventions are unknown. Their identifiability results are up to trivial indeterminacies (permutations and element-wise diffeomorphisms) and identify both the causal graph and the mixing function. Their first theoretical result shows identifiability for two causal variables given one perfect stochastic intervention per node. Their second theoretical result shows identifiability for an arbitrary number of variables when there are two paired perfect stochastic interventions per node. The main text of the paper does not have an experiment section.

Strengths

- This paper frames a problem setting for identifiability that is interesting to causal representation learning, which begins to bridge the gap from existing identifiability results to modern machine learning that occurs on high-dimensional observed data. - They work in a highly general setting where both the causal model and the mixing function are nonparametric, and the targets of the interventions are unknown. In my opinion, the problem framing and the choice of this general setting are the primary contributions of this work even if the theoretical results have limitations. - The paper is generally well-written and well-structured.

Weaknesses

- Their first theoretical result is in a setting with only two causal variables, where you have an observation distribution and one perfect intervention per node. This result would be much stronger in a setting with n>2, as the authors note in the conclusion. - Their second theoretical result is in a setting with arbitrary number of variables, but requires two distinct perfect paired interventions. Requiring these pairs of interventions is not a terribly realistic assumption, even though they don't require the intervention targets to be known. - There is no estimation method or experiment results. Other identifiability papers often contribute an estimation method (e.g. a VAE using a regularizer that encourages sparsity of a mixing function), perform disentanglement experiments in settings that match their theoretical assumptions, or perform ablations on synthetic data where they can control which of their theoretical assumptions are met in order to empirically study the necessity / sufficiency of their assumptions.

Questions

- Does your theory suggest any empirical validation that you could add to this paper? See the last bullet point in "weaknesses" section for empirical approaches that could be relevant to this kind of theoretical contribution.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

- The conclusion section includes a thorough treatment of limitations of this work, which helps future work to extend these results. No concerns about negative societal impacts.

Reviewer E44d2023-08-11

Thank you for addressing my concerns. Please add a version of the above discussions about the assumptions in the revision. I have increased my score :)

Reviewer ZLAN2023-08-11

Response to Rebuttal

Thank you for your response and clarifications. > Clarification regarding assumption (A2’) and (un)known intervention targets. I generally agree with you that your case indeed does not require knowing the intervention targets. My comment was mostly regarding the assumption A2, which currently suggests that this would be the case. By replacing the phrase of "*there exist $n+1$ environments...*" to "*there exist at least $n+1$ environments...*" (line 240), this could be clarified. This is because if you have exactly $n+1$ environments, the intervention targets can be inferred (under some definitions of previous works), and if you have more than $n+1$, the intervention target cannot be inferred anymore. If we take the example of three variables given in the general response, having four datasets, with the first being the observational one, would implicitly give you the intervention targets by numbering the dataset pairs in any arbitrary order. No restriction to isomorphisms of the unknown graph are needed, if one learns an arbitrary graph. If you have more than $n+1$ datasets, this is not possible anymore. > Emperical results It is good to see additional results. Still, the setting and usability of the method is signifciantly limited given the iteration over intervention targets. This requires training up to $n!$ models, a number that can quickly go out of hands for $n>3$ and slightly more expensive datasets to train on. Further, given the small differences and high standard deviations (e.g. diff between 132 and 231 despite inverting the graph), it is unlikely that the empirical method is applicable in challenging datasets at the moment. Thus, the empirical part remains a considerable weakness of the paper. Nonetheless, as mentioned in my original review, the strengths from the theory outweigh this, in particular when the paper is improved by the suggestions of the reviewers.

Authorsrebuttal2023-08-18

Author Response

Thank you for the clarification, continued engagement, and insightful remarks. _____ **Known vs Unknown Targets.** What exactly constitutes known vs unknown intervention targets appears to be a less clearly defined concept in CRL than in the fully observed case. In principle, we agree that *if all possible causal graphs are considered* and exactly $n$ interventional environments with distinct targets are provided, then one can indeed simply call variable $V_k$ the one intervened upon in environment $k$ for $k=1, …, n$, and consider the intervention targets *known in this sense*. In our work, we instead consider a setup in which the causal ordering is fixed to $V_1 \preceq V_2 \preceq … \preceq V_n$ and only graphs consistent with this ordering are considered (see l.167 ff.). This formulation is motivated by starting from the generative process and its fundamental ambiguities (e.g., ordering of the nodes), comes w.l.o.g., and was also adopted in some recent works, see, e.g., Squires et al. [111, Remark 1]. The intervention targets are then considered *unknown w.r.t. this pre-imposed causal ordering*. Our current proofs show that the intervention targets can then be identified from exactly $n$ environments (possibly up to irresolvable partial re-ordering), even if they are not known (w.r.t. the pre-imposed causal order) a priori. ______ **More than $n$ environments.** Regarding the “at least” formulation of allowing $m>n$ interventional environments, this would indeed strengthen the results and break the suggested strategy of assigning target $k$ to interventional environment $k$. In this case, we do *not* know a suitable subset of $n$ environments which contains exactly one intervention for each node. It therefore first needs to be shown that such a subset of environments can be identified. Suppose for a contradiction that we select a subset of $n$ interventional environments which are assumed to correspond to distinct targets in the model $Q$ whereas this is not the case for the ground truth $P$ (i.e., there are actually duplicate and missing interventions). - For the setting of Thm. 4.3 with paired interventions, we can show that this is not possible: Suppose that there are two pairs of environments $(e_a, e_a’)$ and $(e_b, e_b’)$ corresponding to interventions on $V_i$ in $P$, but which are modelled as interventions on distinct nodes $Z_j$ and $Z_k$ in $Q$. Similar to the proof sketch in l.301-303, it can be shown that $V_i$ must then simultaneously be a deterministic function of only $Z_j$ and only $Z_k$. This implies that $\partial \psi_i / \partial z_l =0$ for all $l$ which contradicts invertibility of $\psi$. Hence, only valid subsets will not lead to a contradiction. We can thus allow for any number $m\geq n$ of paired environments (“completely unknown targets”) for Thm. 4.3, as long as there is at least one paired intervention for each node. We will adjust assumption (A2’) to reflect this generalization, and will add a more detailed version of the above argument to the proof. - For the setting of Thm. 4.1 with single interventions, finding a contradiction unfortunately seems more difficult. In short, we can rule out duplicate interventions on root nodes, but for $V_1 \to V_2$ we were not yet able to find a contradiction to selecting a subset of environments corresponding to two interventions on $V_2$. We will continue to investigate this matter, but remark that prior results in simpler parametric settings (e.g., Squires et al. [111], Varici et al. [116]) also require access to a set of exactly $n$ interventional environments, one for each node. We will add a summary paragraph about the subtleties of known vs. unknown intervention targets in CRL to the discussion.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC