Identifiable Shared Component Analysis of Unpaired Multimodal Mixtures

A core task in multi-modal learning is to integrate information from multiple feature spaces (e.g., text and audio), offering modality-invariant essential representations of data. Recent research showed that, classical tools such as {\it canonical correlation analysis} (CCA) provably identify the shared components up to minor ambiguities, when samples in each modality are generated from a linear mixture of shared and private components. Such identifiability results were obtained under the condition that the cross-modality samples are aligned/paired according to their shared information. This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned. A distribution divergence minimization-based loss is proposed, under which a suite of sufficient conditions ensuring identifiability of the shared components are derived. Our conditions are based on cross-modality distribution discrepancy characterization and density-preserving transform removal, which are much milder than existing studies relying on independent component analysis. More relaxed conditions are also provided via adding reasonable structural constraints, motivated by available side information in various applications. The identifiability claims are thoroughly validated using synthetic and real-world data.

Paper

References (77)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer YjhB5/10 · confidence 2/52024-07-08

Summary

This work presents a method for performing a shared component analysis (SCA) in the case of multi-modal unpaired data drawn from a linear mixture. This problem (and method) will be referred to as Unaligned SCA. Unaligned SCA is tackled by matching the probability distributions of the embedded (features) multi-modal data. More specifically, the authors draw inspiration from the traditional adversarial loss used in GANs, and formulate the component analysis algorithm as a min-max optimization problem where a discriminator (a neural network) is trained to maximize the confusion between the alignment of two representations (typically from two different modalities), and then the best alignment is sought in order to “fool” the current discriminator. The alignment matrices are structured so that the algorithm can distinguish between components shared across modalities and private components specific to each modality. The authors claim that the shared components can be identified up to the same ambiguities as those identifiable in the aligned case. Furthermore, the authors explain that while there are other methods that attempt to solve the Unaligned SCA, the conditions of the proposed algorithm are considerably milder. The algorithm is then extended to cases where additional knowledge is present. First the algorithm is modified to accommodate the scenario where the data is generated by a single modality (uniform modality). Then, the algorithm is modified to accommodate the case where some data pairing is available (similar to a weakly supervised case). The authors show that by adding appropriate constraints the shared component can be identified under milder conditions. The author provide first a theoretical analysis with numerical simulations, and then some concrete applications of Unaligned SCA for the problem of domain adaptation (same modality), Multi-lingual Information Retrieval (only in the appendix, same modality) and Single Cell Sequence Analysis (multi modality with and without pairing).

Strengths

- The work tackles an important area of research with potential applications that range from explainability, to SSL, or multi-modal problems. - The method seems very flexible: - The method proposed can work for completely unpaired data but if some paring is available it can take advantage of such additional knowledge. - The method proposed is meant for multi-modal scenarios but it can also work in homogenous use cases. - All parameters are well documented in the appendix and code was part of the additional material.

Weaknesses

I would divide the weaknesses into two groups: the empirical evaluation, of which I am fairly confident about, and the theoretical analysis, of which due to my limited knowledge in this field I am less confident of. I will share here my concerns as solving them might also be helpful for other readers in my position *Empirical Analysis.* I find the empirical analysis weak. The multi-modal and unpaired scenario results are underwhelming, while I find the domain adaptation (same modality) potentially problematic due to the use of CLIP as pre-processing step, and the lack of a strong recent baseline (less problematic than the CLIP reason). - The only practical results where the data are multi-modal and unpaired are the ones shown in Figure 4 first blue dot where the accuracy is about 10%. Since this is the only case presented it is unclear if the method, in practice, cannot cope in this scenario or if the specific problem chosen is particularly challenging (in which case other evidence would be maybe better). - In the domain adaptation experiments the paper reads: “The images are pre-processed by the pretrained CLIP model [34] that uses ViT-L/14 transformer architecture.” It is not clear if ALL baselines used the CLIP embeddings as pre-processing, or if this was only done for the proposed algorithm. For fair comparison the same pre-processing should be applied to all algorithms. - Additionally, CLIP is known to have been pre-trained on a large and diverse dataset and there is a good chance it has been trained on Home-Office and Office-31 too, so it is difficult to appreciate the ability of the proposed method when using such a powerful pre-processing step (which as I mentioned might have been trained on theses datasets, including the their test sets). So making sure the pre-processing is equal is necessary, I would also encourage to present results with a less powerful pre-processing step (e.g., something pre-trained on ImageNet either supervised or SSL style like SimCLR) in order to better distinguish the contribution of CLIP vs every baseline and the proposed model. - The authors use a lot of baselines as comparison however all these baselines seem to be fairly old (all before 2020?). I would suggest comparing with a stronger baseline (either check the leaderboard or here are some suggestions [1-4], note that not all might be immediately applicable). [1] D3GM (https://arxiv.org/pdf/2401.05465) [2] CLUE (https://arxiv.org/abs/2010.08666) [3] LAMDA (https://arxiv.org/abs/2208.06604) [4] SDM (https://arxiv.org/abs/2203.05738) *Theoretical analysis.* I have struggled to follow the theoretical explanation. Specifically, I understand the rationale behind formulation in eq (7) but I would fully agree with it if the samples were paired. I do not understand how 6(b) holds for unpaired samples. I suppose this is the explanation currently presented before assumption 1 but even after reading it I was left with the same question.

Questions

A satisfactory answers to these points could improve the "contribution" (and partially the "soundness") of the work. - Was CLIP used as a pre-process step for all the baselines in the Domain Adaptation? If not, could the authors present those results by using the same pre-processing step? As I mentioned above ideally without using CLIP due to the potential contamination of the test sets. - Could the author explain why 6(b) holds for unpaired samples? Further addressing these less critical aspects could improve the "Presentation" score (and partially the contribution see first point) - I believe the paper would be stronger with more recent baselines as suggested above. - I find unclear how Fig.1 was created. More explanation would help the understanding. - I find Fig 2 not clear: why are there c1 and c2 and only p1, whereas I was expecting p1 and p2 and a common (shared) c? Reading if further it might be that c1 and c2 are the two axes of the common space $c$. If this is the case I’d make sure to clarify it. - Assumption 1 is not clear to me. Is this saying that both points y1 and y2 leave in the same sub-space within the span of $(c, p^{(1)}, p^{(2)})$? If so in which way is this useful? - The authors say the results are an average of 5 runs, which is great, but the standard deviation should also be reported. - I find this sentence confusing: "First, it is unclear if (6b) could disentangle c from p(q). In general, Q(q)x(q) could still be a mixture of c and p(q) yet (6b) still holds (e.g., when both c and p(q) are Gaussian.)" first it says it is unclear if they can be disentangled, but the whole identifiably relies on the ability to disentangle them no? - In a couple of places in the manuscript the authors refer to an experiments where all the results are in the appendix. I would add a brief summary of the results so the main paper is self-standing (this happens in More Synthetic-Data Validation and Application (iii)). - In a few places $p^{(2)}$ appears without a closed bracket, i.e., $p^{(2}$. - Sometimes the authors used the comma (,) to separate thousands but other time they didn’t. I would use a uniform notation.

Rating

5

Confidence

2

Soundness

3

Presentation

3

Contribution

2

Limitations

Yes, the authors have identified and clearly stated three main limitations: - The fact that the conditions presented are sufficient while the necessary are not yet known. - The fact that the method works only for linear mixtures (which limits the expressivity). - Lastly, the fact that the theoretical derivation assume infinite data.

Reviewer YjhB2024-08-09

Thank you for your explanation. I now think I understand where the misunderstanding comes from: by multi-modal I was expecting two completely different data modality (e.g., image and sound). By this definition I thought the only true multi-modal setting was application (ii) with one type of data being RNA sequences and one being ATAC sequences. After a bit of further investigation I think even this setting is not quite multi-modal (as in two separate and different modality) as I believe (but I am not a biologist) both sequences are actually made by the 4 DNA bases. Similarly, the other two applications are mono-modal: application (i) is images (taken from different type of cameras) and application (iii) is text (coming from different languages). I appreciate the in all three applications the data have different distributions, but I think it should be made clear that the focus is not multi-modal settings (as in multi-modality type of data) but rather in same modality with a distribution shift, also known as domain adaptation. I believe changing the title and the explanation to domain adaptation rather than multi-modal would help setting the right expectation for the reader. I would also encourage the authors to provide a short summary of the results of application (iii) rather than delegating everything to the appendix. With all the additional experiments and clarifications provided during this rebuttal I believe the authors have increased the quality of their work to the following: Soundness: 3: Fair -> Good Presentation: 3: Fair -> Good Contribution: 2: Fair The reason for keeping the contribution as Fair is that: - The method was tested only on domain adaptation experiments rather than what I thought was a multi-modality setting. - It was shown in the additional experiments that using a powerful feature extractor (CLIP in this case) is arguably more beneficial than any sophisticated algorithm. I am ok to increase my overall recommendation from 4 to 5.

Authorsrebuttal2024-08-09

We thank the reviewer for their detailed and constructive discussion during the rebuttal. The comment about multi-modality seems to be a terminology issue. We will follow the reviewer's suggestion and change the term to "multi-domain". We will also add the summary of results of application (iii) in the main paper. Nonetheless, the terminology issue doesn’t seem to affect our contributions. We believe that our contribution lies in providing rigorous understanding of the proposed unaligned multi-domain problem structure. Our synthetic and real-data experiments were designed to validate that understanding. The unaligned SCA problem is of great interest as a latent component analysis problem, like ICA, PCA and CCA. Its identifiability has been elusive and our work filled this gap. We wonder if the reviewer could re-assess the contribution from the identifiability research viewpoint, rather than the “multi-modality” vs “mono-modality” viewpoint. In any case, we sincerely thank the reviewer for the comments and discussion, and for pushing us to improve our experiments and presentation.

Reviewer tDhK5/10 · confidence 3/52024-07-13

Summary

The paper considers the identifiability of shared components from a linear mixture. The theory requires multiple domains. However, compared to previous works, the required domains do not need to be aligned in this work. A practical estimation model has been proposed according to the theory.

Strengths

1. The discussion on the related work is comprehensive. 2. The experiments have been conducted on both synthetic and real-world datasets. 3. The proposed algorithm looks pretty neat. 4. Limitations have been discussed in detail together with potential next steps.

Weaknesses

1. Since there are already many works in learning shared components in the nonlinear setting, and some of them can even handle unpaired mixtures, the linear setting appears less appealing in comparison. 2. Assumption 1 is similar to the one used in the previous works, which should be highlighted earlier in the paper. 3. The discussion of the assumption of hyper-rectangle support is missing. Is it restrictive? Maybe some real-world examples could be helpful.

Questions

1. I didn't fully understand the proof in Line 780--does the usage of data processing inequality require a Markov chain? If so, has it been shown in the proof? 2. Could you please elaborate more on the connection between the proposed theory and previous work focusing on identifying content and style variables? It seems like they share a similar goal.

Rating

5

Confidence

3

Soundness

2

Presentation

3

Contribution

2

Limitations

The authors have discussed the limitations.

Reviewer 1Nhs6/10 · confidence 3/52024-07-14

Summary

This work considers a problem similar to classical Canonical Correlation Analysis (CCS), which assumes a linear generative model for data $(x_1, x_2)$: $x_1=W_1z$, $x_2=W_2z$ and aims to identify the underlying components. This problem has been extended previously to include "private information": $x_1=W_1z_1, z_1=[c,p_1]$, $x_2=W_2z_2, z_2 = [c,p_2]$ for common $c$ and independent $p_j$. The current work further assumes that data is *unpaired* and rather than mapping each pair ($x_1$, $x_2$) such that $z_1$, $z_2$ are close together (in some metric), it is proposed that all $x_1$ are mapped to be similar *in distribution* to the mapped $x_2$s.

Strengths

The paper aims to provide rigorous criteria in which underlying generative factors are identifiable in the extended CCA problem it tackles (unpaired CCA with private information). The results show improvement over benchmarks indicating promise to the approach.

Weaknesses

High level: * While I understand the basics I am not an expert in the area of CCA, but I find the paper fairly difficult to follow. More explanation would be helpful, e.g. - [28] why is any linear mixture model ill-posed (is that strictly true in *every* linear mixture case?) - [47] what is meant by "facilitating one-to-many translations", the context/meaning is unclear. * The theoretical part of the paper relates to a simple linear model, but none of the experiments follow this model - e.g. the algorithm is applied to CLIP embeddings, which are not "the data", so to make claims about a simple linear model z=Ax and then apply it to CLIP seems incongruous. This experiments seem to relate more to a CCA-based "loss function" that takes representations and looks to align them/encourage independent factors etc. - other experiments appear to be on discrete data, which the methods doesn't apply to, presumably these are also represented as some intermediate step? - it seems strange to propose a simple linear model, present theoretical results about identifiability that rely on that simplicity but make a dramatic departure in the experiments where the assumptions clearly do not hold and the notion of identifiability is unclear. * if the work does achieve an improvement in a CCA type setting, it seems appropriate to compare with other CCA methods on suitable data. Adding results on more complex data/representations may be of additional interest. Assumption 1 - hard to parse and could be made more clear. - unclear if correctly defined, don't vectors y need to be orthogonal to subspace P? It seems extremely loose to the point of simply saying $P_{c,p_1} \ne P_{c,p_2}$ (specifying where any difference lies to this might add clarity). Theorem 1 - is this saying that if all dims of c are distributed differently, p(z)'s can only match (e.g. under GAN loss) by correctly aligning each dim? If so, that is pretty intuitive and it could be made clearer that you are putting that mathematically and proving it for the sake of rigour. - I have not been through the proof, 5+ pages of proof without a sketch in the paper might be more suitable for a journal as appendices are not typically expected to be reviewed in detail. Overall, there may be useful results in the paper, but in my view it should be re-written to make more clear what it is doing. It seems a confusing mix of simple linear generative model and related CCA methodology mixed with much more complex representations (e.g. CLIP) passed through a GAN + linear layer.

Questions

see weaknesses

Rating

6

Confidence

3

Soundness

2

Presentation

2

Contribution

2

Limitations

see weaknesses

Reviewer 1Nhs2024-08-12

Reviewer response

I realise (as an author) that reviews can sounds attacking. I would like to stress, since your response doesn't seem to acknowledge any change, that I appreciate this line of work and my comments are to improve the paper if possible. * **LMMs & 1-many translations**: these are points of clarity to "the general reader", not just me, I think the paper could be more clear and standalone than it currently is * **Preprocessing**: you mention "linear subspaces", I'm not sure I follow, for sure various representation models give representations that already untangle much complexity in the data (e.g. so that semantically similar items are clustered). CCA acts on the raw data, you are acting on representations. It does not make the approach invalid, but it should be more clearly stated that is what you are doing. You are in effect heavily relying on what other models achieve, which is completely arbitrary with respect to your contribution. In effect you are providing a loss function to wrap around pre-trained representations along the principles of CCA. This is not what is in the abstract for example: "This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned". Given you don't know what the "representation model" has done, it detracts from identifiability claims, which typically refer to the data itself and should at least be caveated. * **Data**: you do run experiments on discrete data: text is discrete. As above you rely on representations that have already done a lot of work in re-representing it. It would be better, in my view, to demonstrate that the linear CCA-type workings actually work as expected on multiple appropriate (simpler) datasets and then show that that still holds for more complex scenarios where non-linear encoders (or similar) have effectively taken the non-linearity into account. Identifiability should relate to factors that those underlying models have identified. * **Thm 1**: I think an intuitive explanation in the paper would improve readability/understandability.

Authorsrebuttal2024-08-13

We would like to stress that we absolutely found the reviewer’s comments valuable for improving the paper’s clarity. It was our fault that we missed adding sentences that commit changes (while concentrating too much on the 6000 character limitation), which was not our intention. We do agree with the reviewer: any suggestion that may help the general readers to better understand the paper is appreciated. We thank the reviewer for the help on clarity and will definitely make revisions accordingly. **[LMM and 1-many translations]** We agree with the reviewer that explanations to these points could make our paper more self-contained. Hence, we will add the explanations (provided in the rebuttal) in a separate "Preliminaries" section in the Appendix with clear pointers in the main paper. **[Preprocessing]** By "linear subspaces", we mean representation spaces where the representations are likely to be linear mixtures of semantic information. As mentioned in our original rebuttal (**[Using CLIP as Pre-processing]**), this has been observed to be the case for embedding spaces of neural networks such as CLIP [R7], word embeddings [R30] etc. We would like to clarify that our method is always applicable wherever CCA is applicable, since CCA shares the same generative model [R18, R29] (also see section 2 Aligned SCA in the manuscript) as ours (the only difference is that CCA further requires the cross-modality samples to be aligned according to their content). Note that CCA also uses pre-processed features representations for complex real-world data (e.g., image, text) [R28, R18]. This is because these complex real-world data might not follow the linear mixture model in Eq. (1) in the manuscript, however the pre-processed representations might. Note that it is common for identifiability works to use pre-processed features for real-data validation of their Theorems [R18]. However, we understand the reviewer’s comment on the applicability to complex data directly. We will explain in more detail in the beginning of the experiment section regarding why preprocessing is involved. We also hope to remark that the sentence in our abstract "*This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned*" is an accurate claim. Note that our claim is for multi-modal **linear** mixtures. Therefore, for complex datasets, it is necessary to find appropriate linear representation spaces. We will add more clarifications/reminders when it comes to the experiment section. **[Data]** We agree with the suggestion of first using simpler raw data and then representations of more complex data to run experiments. The presented experiments in fact may have implicitly reflected this comment. To explain, note that the single-cell data is not pre-processed using any encoder but a normalized (zero mean and unit std) version of the raw-data, which is a simpler dataset as the reviewer mentioned. The more complex image and language data were preprocessed by existing encoders. Following the reviewer’s suggestion on “simpler datasets ---> harder datasets” comment, we will change the order of presenting the single-cell experiment and the other experiments, to make this more explicit. **[Thm 1]** Thank you for your suggestion. We will include a simpler, intuitive explanation of Theorem 1 in the revised version. **References** [R28] Shi et al., 2019. Image Retrieval via Canonical Correlation Analysis. [R29] Ibrahim et al., 2020. Reliable Detection of Unknown Cell-Edge Users via Canonical Correlation Analysis. [R30] Mikolov et al., 2013. Efficient Estimation of Word Representations in Vector Space.

Reviewer 1Nhs2024-08-13

thanks for the response

Thanks, if the proposed changes are made I think the readability and understanding of the paper should improve. I have not been through the proof in detail, but the approach makes sense and the empirical results suggest the method works. Assuming the proposed changes are made, which are not substantive, I recommend the work being published. Score 4 --> 6

Authorsrebuttal2024-08-13

We would like to thank the reviewer for helping us improve the clarity and for the constructive communication.

Reviewer DxMx7/10 · confidence 3/52024-07-23

Summary

This work considers the identifiability of linear latent representations that are shared (i.e., identical) across data modalities, in the special case that they are unaligned/unpaired. The approach leverages GAN-style training to achieve divergence minimization between the latent distribution of each modality. The approach appears to be restricted to two modalities. Under the assumption of shared latents, sufficient (not necessary) conditions for identifiability are presented, which are milder and, thus, more general than existing studies. Further, structural constraints based on side information are introduced to further relax identifiability conditions. Several experiments on simulation and real-world data are provided.

Strengths

Originality : - The work introduces a combination of novel and known ideas in a clever formulation that yields new, less restrictive conditions for identifiability of shared signals from two modalities. - The work differs from and extends previous contributions, dealing two unaligned/unpaired modalities. - The work further relaxes the identifiability conditions via structural constraints based on additional side information that may be available in certain problems. - The manuscript cites related work on identifiability conditions for aligned/paired data, as well as unaligned/unpaired results using stricter ICA conditions, also linking the work to nonlinear studies, adequately indicating the sources of inspiration. Quality : - The work is technically sound, including proofs for identifiability claims. - The theoretical claims are well supported by the experiments. Clarity : - The work is well written and organized, focusing on the key points and contributions. Significance : - The results appear to be quite meaningful, with a potentially wide range of application.

Weaknesses

Quality : - The code is not too friendly to readers, lacking higher 1-to-1 correspondence with the notation in the paper. Suggest improving documentation and tidying up the codebase for readability. - Figure 5 does not seem to replicate well in `synthetic_train.ipynb` --> Clarify - The numerical validations (simulation) are limited to 100,000 samples. Could you illustrate performance at 10,000, at 1,000, and at 100 samples? Many multimodal applications are limited to sample-poor regimes (N < 100), where classical CCA is one of the few performant methods, so it would be useful to assess the performance of unaligned SCA at varying sample sizes, and perhaps include comparable CCA results. Is there a summary measure (like Amari distance) that could be reported alongside figures? Clarity : - The paper contains some typos that limit the clarity. Besides fixing these typos, consider improving the readability of the proofs by being a bit more explicit with "obvious" steps that may be currently omitted. - I think the set L of paired samples was not defined? - Line 59: "transformations identifies" --> "transformations identify" - Line 65: "samples available" --> "samples are available" - Lines 95-96: "the cross-modality samples share the same c are aligned" --> unclear meaning... maybe drop "share the same c"? - Line 117: "to met" --> "to be met" - Line 183: Sentence ends abruptly at: "where \Theta^(q)."

Questions

1. What do you mean by "lift" the constraints in line 131? 2. Although theorem 1(a) does not "require" independence between c and p, isn't that implied/necessary? Otherwise, could you show that Dependence between p and c still yields identifiability of c? Specifically, say, if c1 and c2 are conditionally independent p(c1,c2,p) = p(c1|p)p(c2|p)p(p). 2.a. Does this have to do with Line 149: Q^(q)A^(q) = [Θ^(q), 0] ? 3. Are there obvious limitations wrt differences in the sample size for x^(1) and x^(2)? Are there any provable biases if the data is highly unbalanced? How does unbalanced data affect the identifiability? 4. Is the methodology and identifiability theory limited to the two-dimensional case?

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

- A discussion of the asymptotic computational performance is amiss? Specifically, with respect to d_C and the total sample size. - The paper does not appear to discuss the applicability of the theorems to the case of q > 2 (i.e., more than 2 modalities). - Lacks a discussion of the stability of GAN training, especially with several additional loss augmentations.

Reviewer YjhB2024-08-08

Thanks to these authors for their answers. I appreciate the explanations and the effort in running the domain adaptation experiments using my suggestions. I suggest using the ImageNet features as the main result in the paper and provide the CLIP ones in the Appendix (as it shows that a powerful feature extractor such as CLIP reduces most of the differences among all the algorithms). There is still one aspect I do not fully understand. Could the authors state explicitly which of the three experiment settings (i), (ii) or (iii) shown in the paper are at the same time multi-modal and unpaired?

Authorsrebuttal2024-08-08

(1) We will follow the reviewer’s suggestion and use ImageNet features as the main result in the paper, while moving updated CLIP experiments to the Appendix. (2) All 3 applications (i.e., (i) image domain adaptation, (ii) single-cell sequence alignment, and (iii) multilingual embedding retrieval) are multimodal and unpaired. But their detailed settings vary. We proposed our unaligned SCA approach under three different settings, depending on how much structural information can be exploited. **[Setting 1] Multimodal and unpaired:** The setting uses $\bf{x}^{(i)} = \bf{A}^{(i)} \bf{z}^{(i)}$ where $\bf{z}^{(i)}=(\bf{c},\bf{p}^{(i)})$ for modality $i$. The different $\bf{A}^{(i)}$’s and $\bf{p}^{(i)}$’s both represent the modality discrepancies. Synthetic data was used to validate the setting. We argued the condition needed here for identifiability was too strong for many applications. This was the motivation for us to consider Settings 2-3. (see Sec. 4 Line 203-208). **[Setting 2] Multimodal and unpaired; the modalities share a homogeneous feature space (sometimes called multi-domain setting):** The setting uses $\bf{x}^{(i)} = \bf{A} \bf{z}^{(i)}$ where $\bf{z}^{(i)}=(\bf{c},\bf{p}^{(i)})$ for modality $i$. The modality/domain differences are captured by $\bf{p}^{(i)}$’s. Unlike Setting 1 where $\bf{A}^{(i)}$ varies across $i$, here the mixing systems $\bf{A}^{(i)}=\bf{A}$ for all $i$. This is often called a homogeneous multi-domain setting, which is a more special case of multimodal learning. This setting makes sense when the data $\bf{x}^{(i)}$ for all domains share the same feature space, often used in applications like image-to-image domain adaptation [R17] and image-to-image style translation [R21]. **Applications (i) and (iii) were used to validate the method under this setting**. Application (i) is on domain adaptation of images. The images from different domains are unpaired. Application (iii) is Multilingual retrieval problem. The modalities correspond to unpaired words in different languages. **[Setting 3] Multimodal and largely unpaired (with a small number of paired data):** The setting uses $\bf{x}^{(i)} = \bf{A}^{(i)} \bf{z}^{(i)}$ where $\bf{z}^{(i)}=(\bf{c},\bf{p}^{(i)})$ for modality $i$. The vast majority of data are unpaired. But there exists a small number of paired data (for example, in our experiment of Fig 4, the total number of data in each domain is 1,874. We considered cases where 0 to 256 paired data exist, i.e., 13% at its maximum). This setting is considered realistic in applications such as [R22, R26]. **Application (ii) was tested under this setting**. It corresponds to the single-cell experiment. Modalities correspond to unpaired RNA sequences and ATAC sequences. [R26] *Wang et. al., 2020. Semi-supervised Learning for Few-shot Image-to-Image Translation.*

Reviewer DxMx2024-08-12

Thank you for your clarifications. I have some follow up questions below: Sample size assessment: The Amari distances are quite remarkable (very low) considering the sample sizes. Could you clarify how you estimate \Theta? I recommend reporting this result (e.g., median +/- std Amari values) for every experiment (assuming you can estimate \Theta). CCA result does not need to be reported. Q2: Can you add the note about conditional independence in a footnote? Q3: Can you provide any empirical evidence about how unbalanced data affects identifiability in the different scenarios investigated here? Q4: So the current theory is limited to Q = 2 domains, but can be extended. Could you elaborate further on the point about "benefits" from extra domains? Do you anticipate diminished returns from including extra domains? Complexity analysis: Could you add a discussion about the computational complexity of the proposed model? Do the memory/computation requirements grow linearly/quadratically/other w.r.t. d_C?

Authorsrebuttal2024-08-13

**[Amari Distance Computation]** In our evaluation, we used $\bf{\hat{\Theta}^{(q)}} = \bf{Q}^{(q)} \bf{A}^{(q)}$$(1:d\_C)$, where $\bf{A}^{(q)}$$(1:d\_C)$ represents the first $d\_C$ columns of $\bf{A}^{(q)}$. Note that $\bf{Q}^{(q)}$ is our estimated linear operator, and $\bf{A}^{(q)}$ is the ground-truth mixing system that is available for synthetic data experiments (we only evaluated Amari distance for the synthetic data on our previous reply). Another note is that (thanks to the above discussion) we realized that general matrix distances (such as Euclidean distance) could be a better fit for our case than the Amari distance. To see, recall that we have content identifiability if and only if $ \bf{Q}^{(q)} \bf{A}^{(q)} = [\bf{\Theta}, \bf{0}]$. Hence, we need $\hat{\bf{\Theta}}^{(1)} = \hat{\bf{\Theta}}^{(2)} = \bf{\Theta}$. However, Amari distance is insensitive (invariant) to permutation and scaling, i.e., $\hat{\bf{\Theta} }^{(1)} = \bf{P} \Lambda \hat{\Theta}^{(2)}$ incurs zero Amari distance, where $\bf{P}$ and $\Lambda$ are any permutation and scaling matrices respectively. Hence, we present the Euclidean distance instead of the Amari distance. Additionally, we also report the $\\| \bf{Q}^{(q)} \bf{A}^{(q)}$$(d\_C+1 : d\_C + d^{(q)}\_P ) \\|_F$ which needs to be close to $\bf{0}$ for identifiability. Table 1: Numerical evaluation of identifiability $\\| \widehat{\bf{\Theta} }^{(1)} (1: d\_C) - \widehat{ \bf{\Theta} }^{(2)}(1:d\_C) \\|_{F} .$ |N | SCA | CCA | | :- | :- | :- | |100,000 | 0.009 | 1.368| |10,000 | 0.007 | 1.544| |1,000 | 0.003 | 2.206| |100 | 0.032 | 1.755| |50 | 0.133 | 1.667| |20 | 1.462 | 1.522| &nbsp; Table 2: &nbsp; $1/2 \sum\_{q=1}^2 \\| \widehat{\bf{\Theta}}^{(q)} ( d\_C+1 : d\_C+d^{(q)}\_P) \\|\_{F}$. |N | SCA | CCA | | :- | :- | :- | |100,000 | 0.021 | 0.284| |10,000 | 0.034 | 0.279| |1,000 | 0.002 | 0.329| |100 | 0.043 | 0.368| |50 | 0.131 | 4.092| |20 | 0.747 | 0.755| We will follow the reviewer’s suggestion and add the new experiment (with mean and standard deviation) for the synthetic data experiments in the revised version. **Q2.** Yes. We will add a footnote about conditional independence in the main paper. **Q3.** Thanks for the suggestion. We have run the following experiment with unbalanced data. For the following experiment, the data for two modalities were generated, by sampling the shared component of dimension($D=2$) from VonMises distribution and private components from Gamma and Laplace distributions. The number of samples in the first modality is fixed to 100,000 and the second view ranges from 10,000 to 10 samples. Table 3: Performance of SCA on imbalance data based on following two metrics, **metric1** = &nbsp; $\\| \widehat{\bf{\Theta}}^{(1)} (1: d\_C) - \widehat{ \bf{\Theta}}^{(2)}(1: d\_C) \\|\_{F} $, **metric2** = &nbsp; $1/2 \sum\_{q=1}^2 \\| \widehat{\bf{\Theta}}^{(q)} ( d\_C+1: d\_C+d^{(q)}\_P) \\|\_{F}$. | \# samples in modality 2 | metric1 | metric2 | | :- | :- | :- | |10,000 | 0.008 | 0.025 | |1,000 | 0.025 | 0.015 | |100 | 0.091 | 0.087 | |10 | 1.375 | 0.209 | We will include the above result (with mean and standard deviation) in the revised version. **Q4. [Possible Benefit of $Q \geq 2$ domains]** One foreseeable benefit of having more than two domains is that Assumption 1, when modified for $Q \geq 2$ domains, could be more relaxed. This is because modality variability in general is satisfied if at least two of the total number of domains satisfy the current Assumption 1. Having more modalities can make the chance of Assumption 1 increased. On the other hand, enforcing the distribution matching constraint Eq. (6b) could be more challenging for more than two domains. We will add this discussion in the revised version. **[Complexity Analysis]** The short answer is that **both the memory and computational complexities of the proposed method scales linearly with** $d_C$. The per-iteration computational complexity is $O(B d\_C (d^{(1)} + d^{(2)}) )$, where $B$ is the mini-batch size. The per-iteration memory complexity is $O(B d\_C (d^{(1)} + d^{(2)}) )$ as well. The complexities are based on the fact that we use mini-batch based stochastic gradient-type optimizer. We will add a section in the appendix to detail the complexity calculation.

Reviewer DxMx2024-08-13

Thank you for your responses. Conditioned on the inclusion of all adjustments, edits, results, and discussions the authors have committed to, I've updated my rating on Soundness from Good --> Excellent, largely due to the additional experiments and evidence provided. Keeping my current score rating of 7 (Accept) as my overall impression of the work remains unchanged.

Authorsrebuttal2024-08-13

We thank the reviewer for the constructive comments and discussion. We will make the promised changes as discussed.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC