Response to further points
Thank you for the response and the further comments. We are glad that we are able to make the technical clarifications.
Regarding the remaining concerns about the framing of the contributions, we really appreciate the comments and will try to clarify it via the following changes:
1. Causal disentanglement in the purely observational setting: in the related work section on causal disentanglement (line 81-82), we stated that “our work is the first to establish identifiability guarantees in the purely observational setting without imposing any structural assumptions over the mixing function”. We will modify this claim to “our work establishes identifiability guarantees of causal disentanglement in the purely observational setting, without imposing any structural assumptions over the mixing function”. In addition, we will add a discussion paragraph to clarify that identifiability of latent factors in the purely observational setting has been considered outside of causal disentanglement, such as in the work [2] mentioned by the reviewer.
2. Score-based approaches in causal discovery vs. causal disentanglement: in the related work section on “score matching in causal discovery” (line 91-92), we commented that “extending these ideas to causal disentanglement is difficult, since we do not observe the latent factors and can only estimate the log-likelihood of the observed variables”. We will add in the technical differences of our proof versus the original proof in the causal discovery setting. In particular, we will add a pointer to section 3.2 (which is where our main lemmas and theorems sit), in which we will incorporate our technical clarification in the earlier response: “The proofs in [3] utilize variance properties on the diagonal elements of the Jacobian over the score of the causal variables to derive a topological ordering. While we utilize this result, it is not sufficient given we don't have access to the specified causal variables. In Lemma 1, we prove that we can only estimate this desired Jacobian up to an unknown quadratic parameterized by the matrix , where …”
We hope that these will be sufficient to clarify our contributions and avoid overselling it.
Thank you for the other comments as well. In response to it, we would like to give our perspectives regarding why we believe identifiability up to layers might still be useful albeit its limitations:
in the emerging field of causal disentanglement, full identifiability of latent causal model is not possible without additional assumptions, which is why many works proposed to consider, e.g., structural or sparsity restrictions on the mixing function, or access to single-node interventions. However, as we tried to illustrate in lines 39-44, these assumptions might be limiting and impractical in many settings, which is why we choose to step back and study what can be identified without interventions or structural restrictions. We choose to study nonlinear additive noise models, as they inherit the nice theoretical properties in the causal discovery setting and allow modeling of non-parametric causal mechanisms. In this case, we attain a full theoretical understanding of what can learned, by showing partial identifiability of up to causal layers, which cannot be improved without additional assumptions. Practically, this would mean that more upstream variables in a hierarchical causal structural can be disentangled easier. For example, in the context-style model in [4], our results show that the context variable can be identified up to themselves.
---
References:
[4] Self-supervised learning with data augmentations provably isolates content from style.