On the Parameter Identifiability of Partially Observed Linear Causal Models

Linear causal models are important tools for modeling causal dependencies and yet in practice, only a subset of the variables can be observed. In this paper, we examine the parameter identifiability of these models by investigating whether the edge coefficients can be recovered given the causal structure and partially observed data. Our setting is more general than that of prior research - we allow all variables, including both observed and latent ones, to be flexibly related, and we consider the coefficients of all edges, whereas most existing works focus only on the edges between observed variables. Theoretically, we identify three types of indeterminacy for the parameters in partially observed linear causal models. We then provide graphical conditions that are sufficient for all parameters to be identifiable and show that some of them are provably necessary. Methodologically, we propose a novel likelihood-based parameter estimation method that addresses the variance indeterminacy of latent variables in a specific way and can asymptotically recover the underlying parameters up to trivial indeterminacy. Empirical studies on both synthetic and real-world datasets validate our identifiability theory and the effectiveness of the proposed method in the finite-sample regime. Code: https://github.com/dongxinshuai/scm-identify.

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer kYam7/10 · confidence 4/52024-06-24

Summary

The manuscript proposes novel methods for learning linear structural equation models from partially observed data (i.e., allowing for latent variables). The authors provide graphical identifiability conditions for such models and describe an algorithm for learning structural parameters from data via gradient descent. Results on synthetic and real-world examples suggest that the method works well in practice.

Strengths

The manuscript is clear and well-written. It makes a meaningful contribution to a well-studied topic, which I think will be of great interest to many NeurIPS readers and the causality community more generally. The theoretical sections on identifiability and indeterminacy are well reasoned, with helpful examples along the way. The implementation is elegant and efficient, with compelling results on simulated and real-world data.

Weaknesses

I was a bit confused by the discussion of necessary graphical conditions following Thm. 1. First, I would recommend avoiding terms like "pretty close to...necessary", which doesn't mean much. We get a bit more insight in Remark 1, where it's revealed that condition (i) *is* in fact necessary, but condition (ii) is not. We learn that if (ii) does not hold, then "there are...some rare cases where parameters can be identified." Oddly, we don't see any examples of such structures (either in the main text or the supplement), or learn what "rare" amounts to here (presumably *not* Lebesgue measure zero?) Specific counterexamples would help a great deal! Perhaps some sort of disjunctive graphical condition could do the trick, e.g. condition (i) + [(ii) OR (iii)], where (iii) covers those purportedly "rare" structures that violate (ii) but are still technically identifiable. Alternatively, I would cut all discussion of necessity from this section and move it instead to a discussion section later in the manuscript. (I appreciate that space is tight here, but with an extra page in a final version this could be a more satisfying solution). Minor errata: -All instances of "upto" should be "up to" -It appears that references 46 and 47 are the same -It appears that references 5 and 19 are the same -The word "Gaussian" is occasionally uncapitalized

Questions

-I take it that the identiability conditions of Thm. 1 are untestable? Or perhaps they have some testable consequences in certain settings? If so, this would be very helpful for practitioners! -Would it be possible to perform inference on the learned parameters? For instance, could this method provide standard errors on linear coefficients? -Though I know the method is designed for the partially observable setting, I'm curious how it fares against competitors such as GES or PC when latent variables are absent?

Rating

7

Confidence

4

Soundness

4

Presentation

3

Contribution

4

Limitations

Yes, in Appx. D. May be worth moving this to the main text in the final submission, however.

Reviewer PaWo5/10 · confidence 4/52024-07-08

Summary

This paper investigates the problem of parameter identification in linear causal models, which is important and well-studied task in causality. The authors examine models that explicitly include both observed and latent variables. The identification of parameters in such models has not been studied so far and the authors are the first to formulate this problem and provide the first results in this regard. The main achievements of the paper are the sufficient and necessary conditions for parameter identifiability. Moreover the authors conducted empirical studies on both synthetic and real-world data to validate the proposed methods.

Strengths

Thought causal structure learning in the presence of latent variables has been well studied in the literature, the parameter identification in models that explicitly include both observed and latent variables -- as presented in the submission -- is new. The paper provides non-trivial sufficient and necessary graphical conditions for parameter identifiability and present them in the context of known graphical conditions for structure identifiability provided in [17].

Weaknesses

The sufficient condition for parameter identifiability assumes that G, in addition to conditions (i) and (ii) in Thm. 1, satisfies conditions 1 and 2 presented in section 3.2. I agree that conditions 1 and 2 imply that the structure G can be identified (as shown in [17]) however the sufficient conditions formulated in this way are overall very restrictive. It would be interesting to have sufficient conditions also in the case of structures which do not satisfy conditions 1 and 2: It can happen that G can be identified even if it does not satisfy 1 and 2 or the structure is provided by a researcher / theory. The authors do not discuss if there exist cases which can be identified but which do not satisfy 1 or 2. Also, it is not clear what is the gap between the sufficient and necessary conditions. The next issue is that the authors do not discuss what is the computational complexity of parameter identification in linear causal models, that explicitly include both observed and latent variables. It is not clear to what extent the sufficient and necessary conditions proposed in Section 3 are useful for (numerical) parameter estimation discussed in Section 4.

Questions

Please provide the motivation for considering edges / parameters from observed to latent variables and edges between latent variables. See also my questions above. In Line 69: explain that d=n+m

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

yes

Reviewer Ljpg6/10 · confidence 4/52024-07-12

Summary

This paper investigates the parameter identifiability of partially observed linear causal models, focusing on whether edge coefficients can be recovered given the causal structure and partially observed data. It extends previous research by considering relationships between all variables, both observed and latent, and the coefficients of all edges. The authors identify three types of parameter indeterminacy in these models and provide graphical conditions for the identifiability of all parameters, with some conditions being necessary. A novel likelihood-based parameter estimation method is proposed to address the variance indeterminacy of latent variables, validated through empirical studies on synthetic and real-world datasets, showing the effectiveness of the proposed method in finite samples.

Strengths

The authors consider the problem of parameter identifiability in partially observed causal models with linear structural equations and additive Gaussian noise. This has not been studied in a similar way before in the literature, especially when observed and latent variables are allowed to be flexibly related. Specifically, observed variables are allowed to be parents of unobserved variables. The authors also provide graphical conditions that are sufficient for all parameters to be identifiable and show that some of these conditions are provably necessary. To address the scenario where noise covariance matrix is unknown, they propose a novel likelihood-based parameter estimation method and validate it with empirical studies on synthetic and real-world data.

Weaknesses

The paper considers a restricted setting, allowing only linear relations between observed and unobserved variables and restricting the noise to be Gaussian. The theoretical guarantees hold under the scenario when the noise covariance matrix is unknown. Nonetheless, the authors mitigate this by providing an estimation scheme for the noise covariance matrix from limited data samples. The conditions of identifiability proposed are not fully necessary and sufficient; this is left for future work.

Questions

I have few questions and comments for the Authors: * Line 139: The authors mention that the indeterminacy of group sign is rather minor. if the parameters are identifiable only up to group sign indeterminacy, we still say that the parameters are identifiable. It would be useful to explain why group sign indeterminacy is a minor issue. It might actually depend on the underlying task, and there might be application scenarios where it is not insignificant. Some details on this would be useful in the main paper. * The additive noise is assumed to be Gaussian. Suppose the noise follows another continuous distribution such as Gamma or Student-t distribution. Can we use the identifiability results? If not, what is special about the Gaussian noise here compared to other continuous noise distributions? * Referring to Proposition 1, some matrix inverses need to be calculated to compute the noise covariance matrices. Will the inverses always exist? If yes, can you elaborate on how? If no, how does that impact the application of Proposition 1? * In the experiments section, the authors demonstrate the effect of sample size on the mean squared error (MSE) for the parameter matrix \( F \). It would be useful to also show how accurate the estimates of the noise covariance matrix depending on the sample size used.

Rating

6

Confidence

4

Soundness

3

Presentation

4

Contribution

3

Limitations

The authors clearly describe the problem setting and assumptions. I have mentioned the main limitation of the proposed method in the weaknesses section. I don't think there are any potential negative societal impacts of this work.

Reviewer WKpu7/10 · confidence 3/52024-07-18

Summary

This paper introduces conditions under which DAGs can be recovered in the linear case where some nodes are observed and some are not. This DAG recovery involves computing edge weights between nodes in a causal graph. Nodes are allowed to be latent or observed, with varying types of edge weight indeterminacy depending on the structure of the causal system in question.

Strengths

There is certaintly interest in estimating DAG structure in practice. The work here presents useful results for how and when that estimation may or may not occur given the structure of the causal system. The extension to latent variables is a contribution, although in practice, latent nodes might increase the already difficult task of interpreting inferred DAG structure. The misspecification analysis is useful and the presence of examples in the text helpful (with a few caveats below). I like the title. The framing is at its strongest when the paper articulates the general conditions under which identifiability can and cannot be achieved. As for whether a given causal system in practice meets assumptions for strongest identifiability is in the end, to my eyes, a very difficult question. Overall, the paper is a solid contribution, although I believe it could be improved (see below).

Weaknesses

The text overall is generally well-written, with some caveats listed below. In my reading, the first half of the paper read more clearly than the second half. For example, I couldn't quite piece together from the discussion of estimation whether the estimated graph will be dense (all nodes connected to all other nodes, given the [seemingly?] continuous optimization being done, e.g., Eq 3. If all edges are connected to all other edges, then the relative usefulness seems to be weakened in that usually, investigators seek out a parsimonious representation of a causal system. I was also wondering what more established methods would yield as an empirical baseline (e.g., PC algorithm); currently, the Estimator-LM (estimator with Lagrange multiplers) is articulated as a baseline. This is one of the methods introduced in the paper. An external state-of-art baseline would be most informative. The authors state that "no existing method...can achieve the same goal as ours." If this is because of the latent variable aspect, one could in principle restrict the MSE calculation to edges among observed nodes. In other words, perhaps there isn't a perfect analogue method but an imperfect comparison could be better than none. We could also get a visual comparison of the DAG among observed variables used in Figure 4 from some existing baseline methods for the Appendix. Finally, there is little discussion of uncertainty estimation. Uncertainty estimation in the observed DAG recovery case is hard, even more so here (presumably). There are points where the text could, to me, use more clarity in the discussion: - Condition 1 seems relatively minimal and even intuitive. It seems very hard to know in practice if Condition 2 (line 193) holds or is even a minimal or very restrictive condition. - I appreciate the author(s)' inclusion of Example 2 and Example 3. I think the logic could be made clear, perhaps with additional shadings or labelings that would help us see which sets of nodes and edges are doing what work regarding Condition 1 and Condition 2. - I would make the "pretty close" language in 218 and 255 a bit more precise. Also, starting line 264. I would revise this from, "it has considerable extents of necessity, and could be expected to serve as a stepping stone towards tighter and ultimately the necessary and sufficient condition for the field." to something that also is seomwhat more precise. Also, there is no guarantee that necessity+sufficiency will be found (or perhaps there is an impossibility), so would hedge this possibility somewhat. - Can you define what it means for "QF and F" to "share the same support" in this case? "Support" is often defined in causal inference settings as an event probability falling between 0 and 0; here, F is defined as the matrix embodying the causal edge coefficients, so I believe what is meant here is that QF and F do not share the same set of non-zero entries. Clarifying this would be helpful. I would also consider beginning with the example of indeterminacy before defining it to help the reader see your point intuitively before the formalization. - I would definition 7 into the main text. It is an important definition used multiple times and without it, it is hard to follow the atomic cover discussion. it's also a very short definition. Moreover, I would in general help the reader along by first explaining the concept before formally/technically defining it. Examples: - The term structure identifiability and parameter identifiability should be clearly defined before the terms are used. (I don't think I see a clear definition before use currently; I would move the paragraph beginning on line 171 up in the text, as it is a clear articulate of the point.) In a similar vein, I would explain what atomic covers are going to do before we jump into the definition on line 164. Other details would help this reader: - The MSE up to orthogonal transformation is an interesting metric. Some mention of how this optimization is done would be helpful. - I would add a sentence explaining whether GPU acceleration would or would not be helpful and why. I also noted several minor points listed here: - References used are inconsistent at times. Sometimes, we see reference to "condition (i)", others to Condition 1. - Line 212. Missing space "identified upto the" should read "identified up to the". This same occurs in line 140 ("upto" should read "up to") and in other parts of the text as well. - Line 147. "indetermincay" should read "indeterminacy". This typo occurs a few times in the text. - Line 129. Clarify what is meant by "entails the same observation as that of..." This also appears in line 142 ("Entails the same observation"). I assume this means something about implying the same probability distribution, but helping the uninitiated reader is usually appreciated. - Line 137. "However, if we set f1,2 = 0, then the parameters are not identifiable. These rare cases of parameters are of zero Lebesgue measure so we rule out these cases for the definition of identifiability" -> I would just say "these presumably rare cases of parameters". Probably some justification is needed to articulate why this should be rare in real causal systems. Does any prior literature speak to this? - Line 72. I believe "the causal edge coefficient of the model" should read "the causal edge coefficients of the model". - Capitalize "Gaussian" on line 286.

Questions

- What are the implications of the diagonal covariance matrix $\epsilon_{\mathbf{V}_{\mathcal{G}}}$? - Regarding, "As variables are jointly Gaussian, asymptotically our observation can be summarized as population covariance over observed variable". Wouldn't this statement also apply in finite samples under Gaussianity? - A major benefit seems to be identification of edge weights. If actual interpretation of the edge weights is going to be done in practice, group sign indeterminancy would limit applicability. Would it help to anchor the sign of one edge based on prior science? Guidance? - In Figure 4, are circulate nodes latent and nodes denoted by [LetterNumber] observed? If so, clearly articulate this in the figure label.

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

I see no major negative societal impacts.

Authorsrebuttal2024-08-07

Rebuttal by Authors Part 2

**Q9:** Regarding, "As variables are jointly Gaussian, asymptotically our observation can be summarized as population covariance over observed variable". Wouldn't this statement also apply in finite samples under Gaussianity? **A9:** In finite samples under Gaussianity, our observation can be summarized as the empirical covariance, which is an estimation of the population covariance. **Q10:** If actual interpretation of the edge weights is going to be done in practice, group sign indeterminancy would limit applicability. Would it help to anchor the sign of one edge based on prior science? Guidance? **A10:** Yes, we can always anchor the sign of some edges according to our preference or prior knowledge in order to eliminate such indeterminacy. For example, in Figure 4, if we expect that L2 should be understood as Extraversion instead of non-Exterversion, we can add one additional constraint during our parameter estimation such that the edge coefficient from L2 to E1 ("I am the life of party.") will be positive (as we believe E1 should be positively related to Extraversion). Thank you for your insightful question and we have added a related discussion to our revision. **Q11:** In Figure 4, are circulate nodes latent and nodes denoted by [LetterNumber] observed? If so, clearly articulate this in the figure label **A11:** Yes. We have revised the caption of Figure 4 as you suggested. Regarding the suggestions on writing such as additional shadings for examples, more precise statements for Remark 1, moving definition 7 into the main text and some explanations to the front of corresponding definitions, and correction of some typos, we thank the reviewer and have revised them accordingly. We genuinely appreciate the reviewer's effort and hope that your concerns/questions are addressed.

Reviewer WKpu2024-08-12

Response

Many thanks to the authors for their detailed responses. These answers clarify some of my questions/hesitations; the associated paper revisions should help bolster the contribution too. The discussion of sparsity and optimization will help readers understand the contribution better, as well as the impact of some of the required assumptions. Pondering the question of whether to alter the numerical rating, with the addition of some of the new baselines, I update my view of the paper from, "no major concerns w.r.t. evaluation" to "with good evaluation", and hence move my score to a "7".

Reviewer Ljpg2024-08-08

Re.

Thanks for responding to questions in detail. The clarity of the paper would be improved given that the authors update the paper as they mentioned in their response to my review. I would keep my acceptance decision and score for the paper.

Authorsrebuttal2024-08-12

Author response

Thank you once again for your valuable comments which have helped improve the clarity of our paper. If you have any further questions or insights, we would be more than happy to hear from you. Thank you! Yours sincerely, Authors of submission 11557

Reviewer kYam2024-08-12

Re: Author rebuttal

Many thanks to the authors for their thoughtful replies to my comments. I will maintain my score and look forward to seeing the revised manuscript in the camera ready draft.

Reviewer PaWo2024-08-13

Comments

Dear Authors, thank you for your thorough answers to my questions and comments. It seems the rebuttal addresses all my concerns.

Authorsrebuttal2024-08-13

Dear Reviewer PaWo, Thank you for the positive feedback. We are so happy that all your questions/concerns were properly addressed. Your insightful review comments have helped us further improve the quality and clarity of our paper. We wonder whether you would kindly reconsider your rating based on our responses to your questions. Your consideration is highly appreciated. Many thanks! Your sincerely, Authors of submission 11557

© 2026 NYSGPT2525 LLC