Unpaired Multi-Domain Causal Representation Learning

The goal of causal representation learning is to find a representation of data that consists of causally related latent variables. We consider a setup where one has access to data from multiple domains that potentially share a causal representation. Crucially, observations in different domains are assumed to be unpaired, that is, we only observe the marginal distribution in each domain but not their joint distribution. In this paper, we give sufficient conditions for identifiability of the joint distribution and the shared causal graph in a linear setup. Identifiability holds if we can uniquely recover the joint distribution and the shared causal representation from the marginal distributions in each domain. We transform our identifiability results into a practical method to recover the shared latent causal graph.

Paper

Similar papers

Peer review

Reviewer AWor7/10 · confidence 3/52023-07-01

Summary

The authors study the setting of unsupervised learning where observations belong to several domains, and we only observe the marginal distribution of each domain. A set of latent variables generates the observations, where a subset of latents are shared across domains. The authors provide the first identification results in this setting, assuming that the latents follow a linear SCM, and the observations are an injective linear transformation of the latents. With this model, identifying the (unobserved) joint distribution of the observations equates to identifying the latents. The mapping between the exogenous noises and the observations, as well as the distributions over the exogenous noises, are identified up to signed block permutation. With additional conditions, the authors also identify the causal graph for the latents up to a signed permutation consistent with the topological ordering of the latents. The authors validate their claims with a synthetic data experiment.

Strengths

This is a strong paper that provides the first identifiability results on multiple-domain unsupervised learning where the joint distribution of the observed variables is unobserved. These results are impactful, since this problem setup is well-studied and has practical applications in single-cell biology. Existing approaches are significantly limited by their lack of identifiability, so this paper makes a valuable contribution. The writing style is rigorous, and definitions, assumptions, and results are explained precisely.

Weaknesses

This paper could be improved with more context on how they are extending existing identification results to achieve theirs. Currently, the authors mention which existing results are being used, but do not provide an intuitive description on why they need to be extended, and how they do so. The notation could be improved. Capital letters are used to denote matrices, probability measures, vector- and scalar-valued random variables, sets of nodes, and sets of edges. It would improve readability if you used font styles (e.g. lower-case bold for vectors) to differentiate them.

Questions

This paper makes two contributions. The first is to extend identifiable single-domain linear ICA to the multi-domain setting. The second is to extend single-domain graph identifiability to the multi-domain setting. In both cases, can you intuitively describe what about the multi-domain setting prevents the existing results from holding automatically? For example, if there are no shared latents, the joint can be identified by trivially applying linear ICA separately to each domain. What is the difficulty imposed by a subset of latents being shared, and how is this circumvented?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

The authors made it explicit that this work is primarily about identifiability, and less about scalable algorithms and evaluation on realistic datasets.

Reviewer ny5v6/10 · confidence 3/52023-07-03

Summary

In this paper, the authors address unpaired multi-domain causal representation. In detail, the authors learn the representations of the observed data from different domains that consist of causally. To achieve this, the authors consider the data generation process where the relationship between the latent variables is linear. Based on this generation process, they prove that the joint distribution of observed data and the shared causal structures of latent variables are identifiable.

Strengths

The authors investigate the causal discovery with latent variables from different domains.

Weaknesses

1. There are several works about causal discovery with latent variables under linear and multi-domain case like [1], which also considers the shared causal structure among latent variables and provide identification guarantees. It is suggested that the authors should discuss these works. 2. Moreover, the authors discuss several works about domain translation between unpaired data and claim that none of these works have rigorous identifiability guarantees. However, Multi-domain image generation, image translation, and domain adaptation belong to the proposed setting, and [2][3] have addressed the multi-domain causal representation learning problem recently. And the authors do not consider these works. It is noted that [2][3] considers the multi-domain causal representation learning with nonlinear transformation, which seems to be more general than the proposed setting. 3. As for the identification of joint distributions, it is not clear why the identification of $l, B$, and $P$ can identify the joint distributions of observed data from multi-domains. 4. In section 3, the authors assume that the distribution of errors is non-Gaussian for the identification of linear ICA by not allowing asymmetric distribution. But some distribution like the Laplace distribution is symmetric and they can also satisfy the identification of linear ICA. 5. According to this paper, the authors consider the structure of latent variables to be linear but flexible. In the simulation experiment, the authors only consider three shared latent variables, it is suggested that the authors should consider more latent variables and different structures. 6. Besides, it is suggested that the authors should consider more compared methods and employ other metrics like recall, and precisions to evaluate the performance of causal discovery. [1] Causal Discovery with Multi-Domain LiNGAM for Latent Factors [2] Multi-domain image generation and translation with identifiability guarantees [3] partial disentanglement for domain adaptation

Questions

Please refer weaknesses

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

Please refer weaknesses

Reviewer B2Nu7/10 · confidence 4/52023-07-04

Summary

- The paper considers causal representation learning from unpaired multi-domain data, with latent variables both shared and specific to domains. - Its key contribution is a new identifiability result for linear causal models with non-Gaussian noise, linear mixing function, and a number of other assumptions. - In addition, the authors develop a practical representation learning algorithm for this setting and demonstrate it on toy data. I've read the authors' rebuttal. They have addressed my concerns adequately.

Strengths

- Causal representation learning is an interesting, relevant, and mostly unsolved problem. - The setting considered here (observational unpaired multi-domain data) is practical and well-motivated from single-cell biology applications. - The identifiability result is, to the best of my knowledge, novel, and substantially different from existing results. - As far as I can tell, it is also correct, though due to review overload I have not been able to check the proofs properly. - The paper is extraordinarily well-written. The authors manage to be precise, yet still provide intuitive explanations.

Weaknesses

- The contribution made here has only one real weakness, and that is the host of strong assumptions underlying the identifiability result: 1D causal variables, linear causal model, no causal effects from shared to domain-specific latents, non-symmetric error distributions, different error distributions for each variable, linear mixing function, full-rank mixing function, observed variables include sufficient "partially pure children", and the list goes on. To put it bluntly, this list makes me wonder if this identifiability result present progress on the road to algorithms that work in practice on interesting real-world datasets. - Of course, strong statements such as CRL identifiability require strong inputs, but these need not be in the form of model assumptions, they could also come from the data side. Perhaps it is a bit out of scope for this paper, but I would be curious if the availability of *interventional* data or some other form of auxiliary data would allow the relaxation of some of these model assumptions. - While the authors do a good job in providing an intuition for why these assumptions are needed, I would like to know if there are any real-world problems that satisfy them all. This is partially discussed for single-cell data and a few of these assumptions, but could the authors provide a more complete example that ideally satisfies all assumptions? - There are no experiments to speak of, though I also don't think that all papers need experiments.

Questions

- See above. - Very minor suggestion: there are a few instances of `\citet{}` that should be a `\citep{}`. It speaks for the quality of the writing that I can't think of any other comments here.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

- The paper is very clear about the (many) assumptions in the theory. It also openly acknowledges the limitations of the experiments. - I do not see any particular need for an extensive discussion of societal impacts.

Reviewer aPgx6/10 · confidence 3/52023-07-07

Summary

This work tackles the problem of learning the latent causal structure from multiple unpaired domains. Under a linear non-Gaussian condition, this work presents the identifiability guarantees for the joint distribution over the domains and the causal structure within the shared latent partition. Synthetic data experiments are presented to validate the theory.

Strengths

1. The problem is well-motivated and timely. Unpaired data are prevalent in the wild, and this work provides a rigorous treatment as the initial step to leverage such data in a principled manner. 2. The paper is nicely written, and the theoretical analysis is clearly articulated with sufficient explanations. 3. The theoretical techniques are clearly explained. Connections and attributions to prior work are appropriately introduced, which aids the assessment of this paper’s technical contribution.

Weaknesses

1. The linear assumptions: practical multi-domain (modal) data-generating processes are often highly nonlinear, e.g., images and text. The applicability of the linear assumption may not be as appealing. 2. The heterogeneous noise distributions: pairwise distinct exogenous distributions appear a strong assumption to me and can potentially oversimplify the technical challenge. I would be interested in learning about the necessity of such an assumption.

Questions

I would like to learn about the authors' response to the weaknesses listed above, which may give me a clearer perspective on the paper's contribution.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Please see the weakness section.

Reviewer B2Nu2023-08-10

Thanks for the clear response to my review. In particular, you made a good case for studying linear transformations. I'm less convinced that the other assumptions are often met in real-world situations, though. As a concrete example, you pointed again to single-cell data – thanks, I will look into that. Of course that is a field that is relevant in its own right. I am curious though whether there are any other real-world scenarios to which your results apply? > However, assuming that the distributions of the latent variables are pairwise different and non-symmetric, which is for example satisfied almost certainly by randomly chosen probability distributions on the real line and can thus be considered a weak assumption I would like to politely push against this kind of argument. Of course there are many different distributions on the real line. Yet somehow, in reality, many phenomena tend to follow certain few distributions, in particular normal ones (thanks to the central limit theorem, I guess). I would find concrete examples of real-world systems more convincing than arguments based on some measure that is perhaps not very representative of our world. Thanks again for your comments, I look forward to further discussion.

Authorsrebuttal2023-08-11

We thank the reviewer for their comment and agree that Gaussian distributions appear in many scenarios due to the central limit theorem. However, we want to emphasize again that our conditions for identifiability of the joint distribution are sufficient but, at the same time, also necessary. On a higher level, one might interpret this as follows: If one is willing to assume that conceptually different latent factors also follow a different distribution, then identification of these factors is possible, and otherwise not. Said differently, if the assumptions hold, then our method can be applied, and otherwise no method will do well. Apart from pairwise different distributions, non-symmetry is then required to fully identify the joint distribution whose dependency structure is determined by the shared latent factors. But if one is not willing to make the additional assumption on non-symmetry (for example, due to many Gaussian real-world scenarios), then it is still possible to identify the shared, conceptually different latent factors. This becomes clear by inspecting the proof of Theorem 3.1 and is an important fact that we should add in a remark to the paper. If the error distributions of the latent variables are pairwise different but not necessarily non-symmetric, then the exact joint distribution is in general not identifiable. However, the non-identifiability would only result in sign indeterminacy, that is, the linear effects from the symmetric latents on the domains can be sign-flipped. In terms of other real-world scenarios to which our results apply, we want to point out that unpaired multi-domain data appears in many phenomena apart from single-cell biology. For example, images of similar objects are captured in different environments [1], data from multiple domains is common in large biomedical and neuroimaging datasets [2,3,4,5], or stocks are traded in different markets (data can be downloaded from Yahoo Finance). Under the linearity assumption, our results provide conditions under which a shared causal graph is provably identifiable. It then depends on the specific application to reason about whether or not certain assumptions, such as partial pure children, are justifiable. Moreover, as we explained in our previous answer, we consider our results as a basis for progress on identifiability results in nonlinear setups, such as image data. [1] Recognition in Terra Incognita \ [2] Multimodal population brain imaging in the UK Biobank prospective epidemiological study \ [3] The WU-Minn Human Connectome Project: an overview \ [4] The Cambridge Centre for Ageing and Neuroscience (Cam-CAN) study protocol: a cross-sectional, lifespan, multidisciplinary examination of healthy cognitive ageing \ [5] Training fMRI Classifiers to Discriminate Cognitive States across Multiple Subjects

Reviewer B2Nu2023-08-16

Thank you for yet another clear response, and for your patience with me. I really appreciate the point that you also show non-identifiability of certain settings, and that that's a valuable result as well. I also liked the additional examples you gave. Overall, I am now convinced that the paper makes a valuable contribution to identifiability theory in causal representation learning. It is of a high quality and should be accepted at NeurIPS. I will adapt my score accordingly.

Reviewer aPgx2023-08-13

Many thanks for the thoughtful response. While I still think certain assumptions are a bit unrealistic (e.g., linearity), I gained a better understanding of the pairwise difference assumption, thanks to the response. I will keep my rating as it is.

Reviewer ny5v2023-08-14

Thanks for the response. I find that the authors have discussed [1] and addressed some problems. However, I think that [2] and [3] have addressed the identification of causal representation, for example, Theorem 4.1 in [3] and Lemma 3.1 in [2], which is suitable for the nonlinear scenario. Compared with existing results, I think the contribution might be limited, so I will keep my score.

Authorsrebuttal2023-08-18

Thanks for pointing us to the specific results in [2] and [3]. They are indeed suitable for a nonlinear setup, but we don't think they address learning a *causal* representation in the way we do. In [2], the authors assume a data-generating process in which each observed domain is a function of *shared* content variables $z_C$ and *domain-specific* style variables $z_S^{(i)}$, where both $z_C$ and $z_S^{(i)}$ are latent random vectors. Lemma 3.1 in [2] then shows that the shared latent random vector $z_C$ is block-wise identifiable, i.e., there is an invertible function $h_C$ such that $\hat{z}_C = h_C(z_C)$ for the recovered $\hat{z}_C$. Similar to our case, recovery of the shared latent variables allows for identifiability of the joint distribution. However, the key difference is that we are further interested in the causal relations among the *components* of $z_C$, which are not addressed in [2] but are our central object of study. The authors of [2] only show that the whole vector is identified up to an invertible nonlinear transformation, which does not allow us to draw any conclusions about the causal relations among the individual components of $z_C$. In contrast, we show identifiability of the causal relations (i.e., the graph) between the shared latent variables in Theorem 4.4. The work in [3] is a predecessor of [2], and it shows the same difference in the results obtained.

Reviewer ny5v2023-08-18

Thanks for your response! I have found the difference between these works. I will raise the score.

Reviewer AWor2023-08-17

Thank you for your response, and for providing additional intuition regarding your contributions. I think the assumption of pairwise different error distributions is reasonable. I think this is a strong paper, and maintain my score supporting its acceptance.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC