Thanks for the quick reply!
**Maximal clique**
Sorry about the confusion. The notation $\subsetneq$ means “proper subset”, i.e.if $A \subsetneq B$, then $A$ is a subset of $B$ but not equal to $B$. This is not to be confused with $\not\subset$. Since there is room for confusion here, we will clarify this in the final version.
**Assumption 2**
You’re right that [a] studies a slightly different setting, although the high-level goal of identifying latents under unknown interventions in a measurement model is the same. To clarify, [a] shows the necessity of Assumption 2 under their assumptions, which are slightly different from our setting. We have independently shown that this assumption is needed in our setting with Example 6 in Appendix C.4.
There are three main differences with [a]: (1) They study linear functions while we focus on nonparametric identification; (2) They allow the bipartite graph between latents and observed to be fully connected while we have graphical constraints (Assumptions 1(c) and (d)); (3) They consider noiseless transformations between latents and observed while we allow noisy transformations. Though we briefly touch on similar assumptions in L747 in Appendix C.4, We’ll clarify this further in the paper. Thanks for the suggestion!
**Line 184, footnote 1**
Your intuition is right, however, the situation is more nuanced with latent variables. In Example 7 in Appendix D, we construct two DAGs $G_{(a)}$ and $G_{(b)}$, and two models $P_{(a)}$ and $P_{(b)}$, that generate identical d-separation and CI relations over X. But they differ over (X,H): $P_{(a)}$ satisfies $H_1 \perp H_3 | {H_2, H_4}$, whereas $P_{(b)}$ does not. (This is clear from the extra edge $H_1\to H_3$ in $G_{(b)}$.)
[We realize now that this point was never made explicit, and we will definitely revise this example and the discussion of maximality to reflect this discussion. We’d like to thank you for surfacing this confusion so it can be properly addressed in the final version.]
So, since these models cannot be distinguished on the basis of the observed data P(X), what should we do? We argue that we should only remove an edge if its removal can be justified on the basis of what we actually observe, i.e. the data X. Although $P_{(a)}$ and $P_{(b)}$ can (in principle) be distinguished, we need to observe H to do so, which we cannot do in practice.
More generally, here is what is happening: In general, of course, there are multiple DAGs that are Markov to a given distribution, and the question is how do we decide on the correct “minimal” representation. Without latents, there is no ambiguity: We can always test all possible CI relations and obtain a complete picture to obtain a minimal I-map. With latents, we must be careful:
- Of course, if we can check CI relations over all of (X,H), then the usual notion of a minimal I-map prevails. But in practice, we cannot access P(X,H) since H is unobserved.
- Thus, in practice, we should restrict our attention to information about P(X) only. In this case, we argue that we should only remove an edge if its removal can be justified on the basis of information about P(X) _only_. This is the essence of maximality: We only remove an edge if it follows from the observed data X. Otherwise, we remain agnostic: We do not want to remove an edge that may in fact reflect a “real” dependence over H.
This is the essence of maximality, and the intuition is the same as for maximal ancestral graphs.
As a result, our characterization of the maximal measurement model aligns with the essence of the maximal ancestral graph. Measurement models have latent variables and we only have access to partial information (ie., observed variables). Since two measurement models can encode the same set of conditional independencies over X and as you have pointed out, the absence of edges encodes nontrivial information, the removal of an edge should be justified carefully on the data we have available.
**References**
Thanks for pointing that out! We didn’t realize the author list has changed since we drafted the paper. We will update it accordingly.