Summary
This paper introduces conditions under which DAGs can be recovered in the linear case where some nodes are observed and some are not. This DAG recovery involves computing edge weights between nodes in a causal graph. Nodes are allowed to be latent or observed, with varying types of edge weight indeterminacy depending on the structure of the causal system in question.
Strengths
There is certaintly interest in estimating DAG structure in practice. The work here presents useful results for how and when that estimation may or may not occur given the structure of the causal system. The extension to latent variables is a contribution, although in practice, latent nodes might increase the already difficult task of interpreting inferred DAG structure. The misspecification analysis is useful and the presence of examples in the text helpful (with a few caveats below).
I like the title. The framing is at its strongest when the paper articulates the general conditions under which identifiability can and cannot be achieved. As for whether a given causal system in practice meets assumptions for strongest identifiability is in the end, to my eyes, a very difficult question.
Overall, the paper is a solid contribution, although I believe it could be improved (see below).
Weaknesses
The text overall is generally well-written, with some caveats listed below. In my reading, the first half of the paper read more clearly than the second half. For example, I couldn't quite piece together from the discussion of estimation whether the estimated graph will be dense (all nodes connected to all other nodes, given the [seemingly?] continuous optimization being done, e.g., Eq 3. If all edges are connected to all other edges, then the relative usefulness seems to be weakened in that usually, investigators seek out a parsimonious representation of a causal system.
I was also wondering what more established methods would yield as an empirical baseline (e.g., PC algorithm); currently, the Estimator-LM (estimator with Lagrange multiplers) is articulated as a baseline. This is one of the methods introduced in the paper. An external state-of-art baseline would be most informative. The authors state that "no existing method...can achieve the same goal as ours." If this is because of the latent variable aspect, one could in principle restrict the MSE calculation to edges among observed nodes. In other words, perhaps there isn't a perfect analogue method but an imperfect comparison could be better than none. We could also get a visual comparison of the DAG among observed variables used in Figure 4 from some existing baseline methods for the Appendix.
Finally, there is little discussion of uncertainty estimation. Uncertainty estimation in the observed DAG recovery case is hard, even more so here (presumably).
There are points where the text could, to me, use more clarity in the discussion:
- Condition 1 seems relatively minimal and even intuitive. It seems very hard to know in practice if Condition 2 (line 193) holds or is even a minimal or very restrictive condition.
- I appreciate the author(s)' inclusion of Example 2 and Example 3. I think the logic could be made clear, perhaps with additional shadings or labelings that would help us see which sets of nodes and edges are doing what work regarding Condition 1 and Condition 2.
- I would make the "pretty close" language in 218 and 255 a bit more precise. Also, starting line 264. I would revise this from, "it has considerable extents of necessity, and could be expected to serve as a stepping stone towards tighter and ultimately the necessary and sufficient condition for the field." to something that also is seomwhat more precise. Also, there is no guarantee that necessity+sufficiency will be found (or perhaps there is an impossibility), so would hedge this possibility somewhat.
- Can you define what it means for "QF and F" to "share the same support" in this case? "Support" is often defined in causal inference settings as an event probability falling between 0 and 0; here, F is defined as the matrix embodying the causal edge coefficients, so I believe what is meant here is that QF and F do not share the same set of non-zero entries. Clarifying this would be helpful. I would also consider beginning with the example of indeterminacy before defining it to help the reader see your point intuitively before the formalization.
- I would definition 7 into the main text. It is an important definition used multiple times and without it, it is hard to follow the atomic cover discussion. it's also a very short definition.
Moreover, I would in general help the reader along by first explaining the concept before formally/technically defining it. Examples:
- The term structure identifiability and parameter identifiability should be clearly defined before the terms are used. (I don't think I see a clear definition before use currently; I would move the paragraph beginning on line 171 up in the text, as it is a clear articulate of the point.) In a similar vein, I would explain what atomic covers are going to do before we jump into the definition on line 164.
Other details would help this reader:
- The MSE up to orthogonal transformation is an interesting metric. Some mention of how this optimization is done would be helpful.
- I would add a sentence explaining whether GPU acceleration would or would not be helpful and why.
I also noted several minor points listed here:
- References used are inconsistent at times. Sometimes, we see reference to "condition (i)", others to Condition 1.
- Line 212. Missing space "identified upto the" should read "identified up to the". This same occurs in line 140 ("upto" should read "up to") and in other parts of the text as well.
- Line 147. "indetermincay" should read "indeterminacy". This typo occurs a few times in the text.
- Line 129. Clarify what is meant by "entails the same observation as that of..." This also appears in line 142 ("Entails the same observation"). I assume this means something about implying the same probability distribution, but helping the uninitiated reader is usually appreciated.
- Line 137. "However, if we set f1,2 = 0, then the parameters are not identifiable. These rare cases of parameters are of zero Lebesgue measure so we rule out these cases for the definition of identifiability" -> I would just say "these presumably rare cases of parameters". Probably some justification is needed to articulate why this should be rare in real causal systems. Does any prior literature speak to this?
- Line 72. I believe "the causal edge coefficient of the model" should read "the causal edge coefficients of the model".
- Capitalize "Gaussian" on line 286.
Questions
- What are the implications of the diagonal covariance matrix $\epsilon_{\mathbf{V}_{\mathcal{G}}}$?
- Regarding, "As variables are jointly Gaussian, asymptotically our observation can be summarized as population covariance over observed variable". Wouldn't this statement also apply in finite samples under Gaussianity?
- A major benefit seems to be identification of edge weights. If actual interpretation of the edge weights is going to be done in practice, group sign indeterminancy would limit applicability. Would it help to anchor the sign of one edge based on prior science? Guidance?
- In Figure 4, are circulate nodes latent and nodes denoted by [LetterNumber] observed? If so, clearly articulate this in the figure label.