Linear Causal Representation Learning from Unknown Multi-node Interventions

Despite the multifaceted recent advances in interventional causal representation learning (CRL), they primarily focus on the stylized assumption of single-node interventions. This assumption is not valid in a wide range of applications, and generally, the subset of nodes intervened in an interventional environment is fully unknown. This paper focuses on interventional CRL under unknown multi-node (UMN) interventional environments and establishes the first identifiability results for general latent causal models (parametric or nonparametric) under stochastic interventions (soft or hard) and linear transformation from the latent to observed space. Specifically, it is established that given sufficiently diverse interventional environments, (i) identifiability up to ancestors is possible using only soft interventions, and (ii) perfect identifiability is possible using hard interventions. Remarkably, these guarantees match the best-known results for more restrictive single-node interventions. Furthermore, CRL algorithms are also provided that achieve the identifiability guarantees. A central step in designing these algorithms is establishing the relationships between UMN interventional CRL and score functions associated with the statistical models of different interventional environments. Establishing these relationships also serves as constructive proof of the identifiability guarantees.

Paper

Similar papers

Peer review

Reviewer 8PAL7/10 · confidence 3/52024-07-04

Summary

This paper studies identifiability under unknown muilti-node interventions (soft/hard), with general models (parametrtic/nonparametric) and **linear** mixing functions. This work provides both detailed proof which justifies the main theoretical statement, and a step-by-step algorithm which guides how to achieve identifiability in practice. Overall, I find this work serves as an important step for interventional CRL towards more realistic settings.   ### References [1] Burak Varıcı, Emre Acartürk, Karthikeyan Shanmugam, Abhishek Kumar, and Ali Tajer. Score- based causal representation learning with interventions. arXiv:2301.08230, 2023. [2] Burak Varıcı, Emre Acartürk, Karthikeyan Shanmugam, Abhishek Kumar, and Ali Tajer. Score- based causal representation learning: Linear and general transformations. arXiv:2402.00849, 2024. [3] Julius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele, Armin Kekic ́, Elias Bareinboim, David M Blei, and Bernhard Schölkopf. Nonparametric identifiability of causal rep- resentations from unknown interventions. In Proc. Advances in Neural Information Processing Systems, New Orleans, LA, December 2023.

Strengths

This paper is extremely well written and clearly structured: it communicates clearly motivations, formulation, technical details, and theoretical implications. The experimental results adequately validate the theory in case of a linear causal model.

Weaknesses

1. The proposed UMNI-CRL algorithm is claimed to work with *general* non-parametric causal models; however, the simulation experiment only showed results on *linear* structural equation model. It would be great if the authors could report further experimental results on non-parametric causal models, to align with the theoretical claims. If there is a valid reason why it cannot be done, I am also very happy to hear. 2. Following the previous point, since this approach requires density estimation, it might not be scalable on nonparametric models. But to be fair, this seems to be a common limitation in many interventional CRL works [1, 2, 3]. 3. Linearity assumption on the mixing function is restrictive, but the authors have acknowledged it and discussed possible future directions to overcome this limitation (sec. 6).

Questions

See the first point in **weakness** section. I am very happy to raise my rating if this issue is resolved.

Rating

7

Confidence

3

Soundness

3

Presentation

4

Contribution

3

Limitations

The authors discussed the remaining open problems and limitations in Section 6.

Reviewer NxzW7/10 · confidence 2/52024-07-11

Summary

This paper advances Causal Representation Learning (CRL) by addressing the challenge of using unknown multi-node (UMN) interventions to identify latent causal variables and their structures. The authors develop a score-based CRL algorithm that leverages UMN interventions to guarantee identifiability of latent variables and their causal graphs under both hard and soft interventions, achieving perfect identifiability with hard interventions and identifiability up to ancestors with soft interventions. Their method outperforms existing single-node approaches by ensuring robust recovery of causal structures in more complex, multi-intervention environments.

Strengths

* Extending the causal representation learning to unknown multi-node interventions * Proofs are provided * Pseudocode is provided * Computational complexity is discussed * Limitations are clearly stated

Weaknesses

* The paper primarily focuses on causal models with linear transformations. This limits its applicability in many real scenarios * The applicability of the assumptions in real scenarios was not discussed * The method was not applied on real world-data

Questions

* Can you please elaborate on the computational complexity and on why it is dominated by step 2? * Can you please discuss the applicability of the assumptions in real scenarios? * I think that adding some real world application can increase the impact of this paper. Is it possible to find such an application?

Rating

7

Confidence

2

Soundness

3

Presentation

3

Contribution

3

Limitations

The paper acknowledges certain limitation. One notable limitation is the assumption of linear transformations in the causal models considered. This restricts the applicability to scenarios where causal relationships are adequately approximated by linear relationships. Additionally, while the paper addresses the challenge of UMN interventions, it acknowledges the complexity involved in identifying intervention targets in such settings, which can affect the ability to fully leverage the statistical diversity inherent in UMN interventions.

Reviewer Vtz88/10 · confidence 4/52024-07-12

Summary

This work studies interventional causal representation learning, where one has access to interventional data, to identify latent causal factors and latent DAG in the unknown multi-node interventions regime. The authors consider a setting where the mixing function is linear and the latent causal model is nonparametric. Under the assumption of sufficient interventional diversity, the authors use score function arguments to show that the underlying causal factors of variation (and DAG) can be recovered (1) up to permutation and scaling from stochastic hard interventions and (2) up to ancestors from soft interventions. The authors propose a score-based framework (UMNI-CRL) and evaluate it on synthetic data generated from Erdős–Rényi random graph model.

Strengths

- This work provides significant results in the unknown multi-node intervention setting, which is much more realistic than the common single-node intervention regime. As opposed to other works, this work studies CRL from a more general class of multi-node interventions (stochastic hard and soft). - The paper is well-written, the concepts are explained well, and the theoretical identifiability results add a lot of value to the current CRL literature. - The use of score functions and score differences in the observation space to estimate the unmixing function, especially for the UMN setting, is a novel and interesting approach for CRL. - This work is the first to establish latent DAG recovery in the UMN setting under any type of multi-node intervention for arbitrary nonparametric latent causal models.

Weaknesses

Although the theoretical contribution of this work is strong, the empirical evaluation is quite weak compared to other works in CRL. There are only experiments for n=4 causal variables. There is also no baseline comparison of the proposed framework with other methods in the UMN setting (e.g., [1]). Also, some discussions are a bit abridged and could use more elaboration in the paper (see below for details). [1] Bing et al. “Identifying Linearly-Mixed Causal Representations from Multi-Node Interventions” CLeaR 2024.

Questions

- I would like some clarification on the intervention regularity condition. Specifically, why does the additional term ensure that multi-node interventions have a different effect on different nodes? It would be good to elaborate on this condition when introduced since it is a central assumption that needs to be satisfied for the results to hold. - How do you obtain $\Lambda$ in Eq. (14)? It seems that this matrix encodes the summands with the latent space score differences. However, since the distribution of the latents is unknown, how would you go about estimating $\Lambda$ and score differences $\Delta S_X$ in general cases of nonparametric distributions? - How do you learn the integer-valued vectors $\mathbf{w}$ in Stage 2 of the algorithm? From Eq (18), it seems that $\mathcal{W}$ is a fixed predefined set and you choose the vectors $\mathbf{w} \in \mathcal{W}$ that satisfy a specific condition in the algorithm. To my understanding, this is central to recovering the approximate unmixing $\mathbf{H}^*$ up to a combination of the rows of the true unmixing $\mathbf{G}^{\dagger}$. I would appreciate it if the authors could elaborate on how this procedure was done. - From Appendix A.8, it seems that $\kappa$ is determined by the number of causal variables $n$. Could the authors give some more intuition on what $\kappa$ represents in Stage 2 with respect to how the unmixing is recovered? - Are there any distributional assumptions on the exogenous noise in the latent additive noise causal model? - It seems that the UMN hard intervention result (Theorem 1) requires a latent model with additive noise. Would perfect recovery still be possible for latent models with non-additive noise under UMN hard interventions? - The empirical results suggest that increasing sample size improves DAG recovery, which is intuitive. However, what do the results look like as the number of causal variables scales up? Currently, the authors only show results for n=4 latent causal variables. I only offer this as a suggestion due to the short rebuttal period. - How would the assumptions made need to change to be applied to general mixing functions? I know that generality in one aspect of the model (i.e., general SCM) may require other aspects to take some parametric form (i.e., linear mixing) for identifiability guarantees, but do the authors have any intuition on how to achieve identifiability results for the UMN setting in a completely nonparametric setup?

Rating

8

Confidence

4

Soundness

4

Presentation

3

Contribution

3

Limitations

Limitations are discussed in Section 6.

Reviewer q4pN6/10 · confidence 4/52024-07-12

Summary

This paper extends previous results on using score function for causal representation learning to the settings with unknown multi-node interventions. This new setting poses significant new challenges as opposed to the single node intervention case. The author first present theoretical identifiability result on hard interventions with latent additive noise model and on soft interventions. They then propose an algorithm called (UMNI)-CRL and test it on synthetic linear Gaussian dataset.

Strengths

The paper is clearly written, easy to follow and with good motivations.

Weaknesses

1. The transformation from latent to observed is noiseless, which could be a limitation. 2. Line 199 says that: “This regularity condition ensures that the effect of a multi-node intervention is not the same on different nodes”. But how realistic or neccessary is this condition? It seems like it is very possible that an intervention can cause two downstream nodes to have the same effect although these two nodes is not influenced the same by all type of interventions. 3. The experiments are only on synthetic dataset but I don’t think that is a big issue. 4. Some potential missing citations [1] Kumar, Abhinav, and Gaurav Sinha. "Disentangling mixtures of unknown causal interventions." *Uncertainty in Artificial Intelligence*. PMLR, 2021. [2] Jiang, Yibo, and Bryon Aragam. "Learning nonparametric latent causal graphs with unknown interventions." *Advances in Neural Information Processing Systems* 36 (2024).

Questions

1. (UMNI)-CRL requires estimating the score function. How do you ensure a good estimate of the score function to unsure that the algorithm is useful in practice? 2. One small question: on line 141-143, it is mentioned that if a node is not intervened on, perfect identifiability is not possible. But there are cases like A→B where I don’t need to intervene on A?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

1. Theorem 1 only works for additive noise model. 2. The transformation from latent to observed noiseless. 3. Experiments are only on synthetic dataset. 4. I am unsure if the algorithm is practical because it needs to estimate the score function.

Reviewer GLEY7/10 · confidence 3/52024-07-15

Summary

This paper introduces new identifiability results for CRL in environments with unknown multi-node interventions. It shows that, with sufficiently diverse interventional environments, one can achieve identifiability up to ancestors using soft interventions and perfect identifiability using hard interventions. The paper also provides an algorithm with identifiability guarantees.

Strengths

- The paper tackles the complex and underexplored multi-node intervention setting. The established identifiability can be crucial for extending current CRL theories into more practical contexts. - The introduced algorithm that leverages score functions with different interventional environments is also interesting and insightful. - The paper is well-motivated and articulated with high clarity.

Weaknesses

- The proposed algorithm, while theoretically sound, seems computationally demanding. In fact, even a 4-node low-dimensional case requires a large number of environments and samples. The paper could benefit from a deeper discussion on the scalability of the algorithm. - The current evaluation of the algorithm is limited to synthetic simulations. Expanding it to more realistic datasets would substantively improve its practical significance.

Questions

How effectively does the proposed algorithm scale to more nodes and higher dimensions?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The paper acknowledges its main limitations in the reliance on linear transformations.

Reviewer 8PAL2024-08-08

I thank the authors for clarifying and providing additional experiment results. I increased the score correspondingly.

Reviewer Vtz82024-08-09

I greatly appreciate the authors taking the time to answer my questions and provide clarifications. My questions and concerns have been addressed quite well in the response. The new empirical results for a larger number of causal variables further strengthen the paper. I believe this is a high-quality submission with significant theoretical results of great interest to the CRL community. Thus, I raise my score to 8.

Reviewer NxzW2024-08-11

I thank the authors for the response. After reviewing the reviews and considering the responses, I will raise my score to 7.

Reviewer GLEY2024-08-13

Thank you for your response to my question. I increased the score to 7.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC