Summary
The paper discusses the task of identifying causal variables from high-dimensional observations under non-parametric mixing functions and causal mechanisms. This is done under the assumption of single-node, perfect interventions being available for all causal variables, as well as distinct paired perfect interventions in the case of having more more than two causal variables. The paper proves that causal variables are identifiable under this setup, when taking additional assumptions on the interventions being sufficiently different from the observational distribution. Thereby, weaker assumptions are possible for 2 variables than for more variables. Finally, the paper sketches possible implementations of learning algorithms for this setting.
Strengths
The paper is overall well written, even if it is aimed at researchers in identifiability and/or causal representation learning (CRL) specifically. A consistent notation is used throughout the paper, and all assumptions are clearly stated before the theorems. It is appreciated that proof sketches have been included in the main paper to support the claimed theorems and make the main paper a bit more standalone. The paper discusses all necessary related work and puts itself into context of the current field of research.
The main contribution of the paper is its theoretical result. It extends the domain of identifiable causal representation by considering yet another setup, where environment pairs with single-node, perfect interventions are given. The benefit of this setup is that it does not require counterfactual observations, while supporting a large function class despite taking needed assumptions on the interventions. The proofs for supporting the claimed theorems are given in the appendix, following common proof strategies in CRL. The proofs appear sound and intuitive, although a very careful check of the proofs was not possible during the review period. Overall, it is a good contribution to the theoretical identifiability in CRL.
Weaknesses
While the derived theory in the paper puts weaker constraints on the mechanisms of the causal variables and mixing function, its assumption of having access to single-node, perfect interventions on all causal variables is restrictive. Being able to perform an intervention on a variable is already commonly considered expensive or often not easily feasible, especially if it is a perfect intervention and single-node. However, doing this twice and even different between the two setups is challenging. Further, obtaining such a dataset requires non-trivial prior knowledge of the causal system, since it necessitates the ability to perform such single-node, perfect interventions on causal variables that are yet to be identified. The paper misses to give real-world examples to motivate the setup and its assumptions, which puts it in a more limited spot.
Besides the theoretical results, it is also important to validate the setup and the practicality of the theory in empirical studies. The paper only sketches some potential ideas, where all unknown parts are learned. However, optimizing the latent encoder, the causal graph, and the intervention targets all at the same time is not trivial as shown in previous works. Further, the appendix shows some limited results on a generative model, where one would need to iterate over all possible causal graphs and intervention targets. Still, this is not practical for systems larger than very few causal variables or high-dimensional observations.
The paper states that the intervention targets are not known. However, under the identifiability up to permutation, the intervention targets in this setup appear to be known. Specifically, assumption (A2') states that there exist $n$ environment pairs, where each pair intervenes on a different causal variable. Thus, the intervention targets for these pairs, as stated in the assumption, are known as $\pi(i)$. Since the variables cannot be identified up to permutation $\pi$ anyway, permuting the causal variables and thus the targets are still considered to be the same targets in the same identifiability class, e.g. as in the works cited for known intervention targets [69, 70]. Thus, the claim of unknown intervention targets appears not valid given the assumptions, or the assumptions should be clarified to e.g. have at least $n$/$n+1$ environments.
### Typos:
- Table 1: 'Causal Representation Learning'
Questions
### Review summary
The theoretical results of the paper provide a new setting under which causal variables are identifiable in CRL. However, the paper is limited by its strong reliance on single-node, perfect interventions and very limited empirical study. I consider the theoretical results outweighing the drawbacks a bit, although the paper would strongly benefit from empirical validation of the setup. Thus, my recommendation is 'Weak Accept'.
### Questions
- What is a real-world scenario in which the setup of pairs of single-node, perfect interventions is practical and common?
- Do you require the knowledge of the intervention targets up to permutation, or do you allow for more environments/environment pairs as long as each variable has been intervened upon once?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
Limitations have been discussed in different parts of the paper.