Summary
Causal analogs to de Finetti's theorem are proven, showing that if in an exchangeable distribution certain conditional independences hold, the distribution can be seen as being generated by a DAG with a latent variable corresponding to each node, determining that node's causal mechanism. It is further shown that knowing these conditional independences allows unique identification of the DAG. A causal discovery algorithm is presented leveraging these results, and evaluated in a synthetic data setting where data come from very many different environments.
Strengths
* The causal de Finetti's theorems are philosophically interesting, similar to how the original de Finetti's theorem can play a role in the justification of Bayesian inference. The resulting graphical models, expressing disentangled causal mechanisms in terms of latent variables (Figure 1b, right), are very insightful and deserve to be commonly known in the community.
* The resulting conditional independences, which are testable if data are available from multiple environments, are a very useful ingredient for causal discovery algorithms in such settings.
Weaknesses
* The related work section only considers causal discovery, not the causal de Finetti theorem itself. Such references should also be listed, because the paper is claiming a contribution in this area. For instance, [Dawid 2021] also uses exchangeability to do causal inference.
* Also about the related work section: the sentence "These algorithms on non-i.i.d. grouped data all demonstrate success, though it is unclear why grouped data enable causal structure identification." - I disagree strongly with this, there is a very good understanding of why multiple environments (possibly including interventional data) help with causal discovery, in the papers you list as well as in the causal inference literature as a whole.
* Implications of the causal de Finetti theorem are listed without sufficiently arguments, and I believe overstated. See question about line 139 below.
* Here is a counterexample to Lemma 2: $X_i \rightarrow X_a \rightarrow X_b \leftarrow X_c \leftarrow X_d \rightarrow X_j$ with $X_b \rightarrow X_j$. Then $X_d \in S_n$, so only $X_b$ is conditioned on, but this opens the listed path.
(Separate from this, in line 294 explaining the lemma, I think "non-directed" should be "open": also directed paths other than the 1-arrow path should be blocked.)
* Experiment (some things are unclear to me now, see questions below):
* other methods are not really fit for this scenario
* 1 x-y pair per environment (?): this disconnects the experiment from the theory
**References:**
[Dawid 2004]: Probability, Causality and the Empirical World: A Bayes-de Finetti-Popper-Borel Synthesis, Statistical Science , Feb., 2004, Vol. 19, No. 1 (Feb., 2004), pp. 44-57
[Dawid 2021]: Decision-theoretic foundations for statistical causality, Journal of Causal Inference 2021; 9: 39-77
Questions
* Line 139: "one can separately manipulate each latent variable controlling different mechanisms" - Does this follow from equation 5? How? ([Dawid 2004] warns that de Finetti's theorem only establishes that the latent variable exists in our minds, not in the real world.)
* In the experiment, I assume $\tilde{N}^e$ is a 2-element vector? What about $N^e$? And how many x-y pairs are sampled per environment?
Suggestions to improve the language (not relevant for my assessment of the paper):
* the spelling of "i.i.d." is inconsistent (sometimes with spaces, sometimes with the final . missing)
* line 92 & 595: "infinite exchangeable" -> "infinitely exchangeable"
* the final sentence of section 2.2 doesn't parse ("due to"&"that underlies"; "involving observations are i.i.d.")
* below definition 2: "does not hold for all" -> "does not hold for any"
* line 140 & 141: "supporting mechanisms" -> "supporting that mechanisms" (2x)
* line 147: "one implicitly make" -> "makes". Similar in appendix A.
* Theorem 3 & 4: I'd replace "The sequence is" by "If", add "the sequence is" to point 1, and remove "if" from point 2
* line 201: "decide" -> "deciding"
* line 220: "process" -> "processes" (also elsewhere)
* Definition 5: "Given $P$ is" -> "Let $P$ be"; "Given $G$ be" -> "Let $G$ be"; "a ADMG" -> "an ADMG" (this also appears in Def 6); "read-off" -> "read off"
* Definition 7: There is only one mapping fitting this definition, so define it straightaway instead of defining when something "is an ICM operator". Line 242 has a double "other"
* line 280: $<$ -> $\leq$
* line 287: "topological ordering" is not the right concept here, as that is a total order
* line 629: "exchangeable" -> "exchangeability"
* line 679: "$l < k + 1$ -> $l \leq k$
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
Yes (except as discussed above)