Generalizing Nonlinear ICA Beyond Structural Sparsity

Nonlinear independent component analysis (ICA) aims to uncover the true latent sources from their observable nonlinear mixtures. Despite its significance, the identifiability of nonlinear ICA is known to be impossible without additional assumptions. Recent advances have proposed conditions on the connective structure from sources to observed variables, known as Structural Sparsity, to achieve identifiability in an unsupervised manner. However, the sparsity constraint may not hold universally for all sources in practice. Furthermore, the assumptions of bijectivity of the mixing process and independence among all sources, which arise from the setting of ICA, may also be violated in many real-world scenarios. To address these limitations and generalize nonlinear ICA, we propose a set of new identifiability results in the general settings of undercompleteness, partial sparsity and source dependence, and flexible grouping structures. Specifically, we prove identifiability when there are more observed variables than sources (undercomplete), and when certain sparsity and/or source independence assumptions are not met for some changing sources. Moreover, we show that even in cases with flexible grouping structures (e.g., part of the sources can be divided into irreducible independent groups with various sizes), appropriate identifiability results can also be established. Theoretical claims are supported empirically on both synthetic and real-world datasets.

Paper

References (44)

Scroll for more · 32 remaining

Similar papers

Peer review

Reviewer qJ697/10 · confidence 3/52023-07-05

Summary

This paper utilizes the structural sparsity assumption on the support of the Jacobian matrix of the mixing function to extend the identifiability of nonlinear ICA in more settings including under-completeness, partial sparsity and source dependence, flexible grouping structures. It is a technically solid paper supported by theorems, proofs and experiments.

Strengths

1. Overall, the manuscript is well written with clear organization, comprehensive literature review, technically solid theorems, detailed proofs and promising experiment results. 2. This work addressed some limitations of theorems about identifiability with Structural Sparsity in Zheng et al. 2022 and extended nonlinear ICA with Structural Sparsity to more general settings. The proposed theorems could be more practically useful in real-world datasets. 3. The notations, theorems and proofs are clear in general.

Weaknesses

1. This work is interesting, and it would be great if code is provided to replicate the results. Please consider making the code publicly available. 2. The meanings of some notations are not clear. See Questions. 3. Ablation study. The author(s) only evaluated MCCs w.r.t. the number of sources. In Figure 3, the MCCs for 8 or 10 sources are a bit low so I wonder if more samples can help to improve MCCs. It would be more informative and convincing if more experiment configurations are considered (e.g., number of samples, various grouping structures) to demonstrate the effectiveness of proposed Theorems. 4. Ablation study. Though the result comparison seems obvious visually, the authors should consider performing statistical tests to compare results between proposed methods and baseline method. 5. Minor: "exits" should be "exists" at line 236.

Questions

1. Theorem 3.1: Is $|\mathcal{F}_{i,:}|$ the $L_0$ or $L_1$ norm of $\mathcal{F}$? Is $\mathcal{C}_k$ a minimal set of sample indices to uniquely identify source $k$? How does the assumption ii show Structural Sparsity? The regularization constraint $|\hat{\mathcal{F}}| \leq |\mathcal{F}|$, which induces sparsity, should be included in Theorems. Also, I note that the author(s) tried to explain the assumptions in the following paragraphs, but I would suggest to describe the Theorems, at least Theorem 3.1, in plain words so that readers can better understand the Theorems. Or at least explain the notations (e.g., $\mathcal{C}_k$) which are not explained in Section 2 Preliminaries. 2. Theorem 3.1: Zheng et al. 2022 also proposed a Theorem on the undercomplete case. Could you kindly clarify the novelty between your proposed Theorem and that proposed in Zheng et al. 2022? 3. Theorem 4.1: The author(s) claimed that we do not need to know the dependence structures or the number of dependent sources, but it is not intuitive to me how Theorems 4.1 and 4.2 uncover the dependence structures and the number of dependent sources. Could you please clarify? 4. Theorem 4.2: What are $u_1$ and $u_2$? Are they two different sets of auxiliary labels? 5. Lines 281 - 283: The claim on multi-modal data is unclear. Could you please clarify how to identify linkage across multiple modalities?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer m3QW6/10 · confidence 2/52023-07-05

Summary

The article serves as an extension to the work of Zheng et al. 2022, which posited the identifiability of nonlinear ICA based on specific structural sparsity assumptions related to the mapping of sources and mixtures. This current article expands on that by addressing the undercomplete case—where the number of mixtures exceeds the number of sources—and furthermore, it relaxes both the sparsity and source independence assumptions to yield more general identifiability results. In the end authors provide some numerical examples to illustrate the applications of their identifiability theorems.

Strengths

The article tackles the foundational issue of the nonlinear inverse problem, making significant assertions regarding fundamental identifiability theorems. These are premised on assumptions of partial independence and structural sparsity.The problem under investigation is a fundamental problem and the article offers some important results for this problem.

Weaknesses

The most significant shortcoming of the article lies in its presentation, particularly in its explanation of the core assumptions underpinning the theorems, as well as the motivation for the conditions applied in these assumptions. Absent a solid grasp of these theorems, it becomes challenging to properly evaluate the paper's contributions and ascertain its potential impact. In terms of specific issues: * Concerning Theorem 3.1: This theorem seems to aim at generalizing Theorem 1 from Zheng et al., 2022 for a complete (m=n) case to an undercomplete (m>n) case. The identifiability results for ICA setups typically do not depend on a specific estimator choice or estimation algorithm. However, the statement of Theorem 3.1 seems rather unclear in this context. According to the article's notation, \hat{f} refers to a specific estimate of the mixing function. Both the set of support matrices \mathcal{T} and the support \hat{\mathcal{F}} depend on this particular estimate. The assumption (i) used in Theorem 3.1 is based on \mathcal{T} and \hat{\mathcal{F}}, hence, this condition appears to be linked to a specific choice of the estimator for mixing. It would be beneficial if the authors could clarify whether the assumption (i) must hold on a particular \hat{f}, a specific set of functions, etc. This clarification will likely impact the proof in the supplementary material and the explanation given between lines 129-136. * Line 97: Traditional or linear ICA does not necessarily require m=n. * Line 114: Should \mathcal{S} be \mathcal{A}? * Line 111 vs Line 536: The symbol \mathcal{T} has two differing definitions - the set of matrices sharing the same support as T(s) and the support of T(s) itself. * Theorem 4.1: The vectors 'w' - whose independence implies identifiability - have not been sufficiently motivated or explained. * Similar comments can be made for other identifiability theorems.

Questions

* Line 190: Could you clarify what is meant by "changing" sources? * Line 195-196: 6, the phrase "For sources s_D, they do not need to be mutually independent as long as they are dependent on the variable u" is somewhat confusing. Subsequent equation (3) implies that S_D and S_I are conditionally independent when conditioned on u, and that the components of s_I are independent. What is the necessity for this latent variable u? Perhaps the authors could shed some light on this. * How does Theorem 4.1’s contribution compare with those from Khemakhem et al. (2020a) and Sorenson et al. (2020)? * Lines 215-217: Could you elucidate what condition (i) in Theorem 4.2 represents? What do s_d and B_{s_I} signify? Perhaps the discussions on lines 229-247 should precede Theorem 4.2 to better contextualize its contents and results, potentially with more lucid explanations. * Regarding Figure 4 and 5, could you expound on how these examples pertain to the nonlinear ICA setup (what are the sources, which appear to be images, and what are the nonlinear mixings)? How are the interpretations in the captions of these figures derived? Could you also elucidate how these examples relate to the identifiability theorems presented in the article and the conditions stipulated in these theorems?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

3 good

Limitations

Yes, the authors adequately addressed the limitations.

Reviewer nmMP6/10 · confidence 3/52023-07-07

Summary

The paper extends identifiability theory of nonlinear ICA (NICA), and deep latent variable models in general, by utilizing structural sparsity. In particular, previous works have shown that NICA can be identified if there is some observed auxiliary data or latent dependencies that essentially capture the inductive biases in the data generative process. The approach of structural sparsity (Zheng, '22) instead takes an alternative approach, namely constraining the nonlinear mixing function and its Jacobian. In this paper the authors extend that work as follows i.e. assuming structural sparsity and some additional assumptions: 1. the authors show identifiability for a situation in which the mixing function is injective rather than bijective i.e. undercomplete case / i.e. smaller latent dimension than observed dimension. This is important for e.g learning low dimensional, semantically meaningful, interpretable latent features 2. the authors show identifiability for a situation in which the structural sparsity principle applies only to some of the independent components 3. further identifiability is shown for situation where not all the latent components are independent, rather components form independent subspaces. In fact some components may be dependent, conditionally independent or have some grouping structures.

Strengths

First, this paper is in general of good quality in that it is well organized and, in general, clearly written. Main strengths: a.) most significantly, authors remove several strong limitations of previous works and extend identifiability of structural sparsity to undercomplete and case where not all latent components are structurally sparse (in those situations the remaining sources are shown to be identifiable). As a result these ideas are now more applicable to realistic data and scenarios. b.) These results have been reached, mostly, without too strong additional assumptions. For instance, it is shown that the necessary assumptions are more likely to hold in this new undercomplete case which is encouraging! c.) authors bridge gap between structural sparsity and the previous works that assume auxiliary variables. in particular, this work allows unconditionally independent components to follow structural sparsity and components which are conditionally independent given auxiliary variables. Whilst arguably to be expected, it is important to show this result (but see below for potential related weakness)

Weaknesses

In general the weakness of this paper is that provides only few theoretical advances (albeit important; as mentioned above) but provides little beyond that. In particular, this is the results of: 1. Contribution of the paper is not as significant as the authors describe, or at least there is limited coverage of relevant works 2. Potential problems in some of the identifiability theorems 3. Novelty is limited to identifiabiltiy theorems -- no new algorithms 4. Experiments are lacking I will expand on each of these points below: More detail for 1.): In particular the authors state that "Therefore, we establish, to the best of our knowledge, one of the first general frameworks for uncovering latent variables with appropriate identifiability guarantees in a principled manner". I think this is too vague and general and fails to acknowledge the generality of some other works -- your work can be novel whilst admitting the generality of some other works too. First, Kivva '22 show a very general framework for identifiability by making, arguably, less strong assumptions on the mixing functions -- currently the work of Kivva '22 is only mentioned later on in section 3 (and even there in a problematic manner as ill point out below). Due to the generality of the results in Kivva, I would expect their result to be discussed in the introduction / early on in the text and tell the reader why yours is better or at least different. For example Kivva make different type of assumptions on the mixing function (piecewise affine) etc, while you on the sparsity. Second the work of Halva '21 (disentangling identifiable features) provides another very general framework and unlike what you claim, it is not limited to time-series but to any dependencies of arbitrary order, and also does not require condition independence on some auxiliary variables but rather also assumes unconditional independence. More detail for 2.): In Theorems 4.1 and 4.3 $S_d$ is "identified up to an invertible transformation". Surely if something is identified just up to invertible (vector-valued) function then we are not doing any better than nonlinear ICA i.e. we are essentially where one started and thus we have not identified anything. To me this is misleading and not a publishable identifiability result (if I have understood correctly -- please correct me if I'm wrong and I'll adjust my score accordingly). Authors do acknowledge this point but rather than talking about it they make a vague remark that "Thm. 4.1 may be helpful for some tasks that do not necessitate the recovery of each individual source, such as domain adaptation." This does not suffice in my opinion. And similarly about Thm 4.3 they say: "there exists an invertible transformation $h_{c_i}$ which is analogous to the previous element-wise indeterminacy. Consequently, even when dealing with mixtures of high and one-dimensional sources, like in the case of multi-modal data, we can still recover the hidden generating process to some extent." Again I think this is bit generous and hiding the fact that $\mathbf{s}_{c_i}$ is fully unidentifiable in the sense of nonlinear ICA. At least this limitation must be admitted more clearly -- preferably its usefulness would be shown empirically. More detail for 3.): There is a simple regularization term of the jacobian added -- but this is heurestic (vs. mle methods) from previous work. Undoubtedly there could be work done towards what is the best way to estimate a model that assumes structural sparsity but such is not done here. More detail for 4.) An important question is whether structural sparsity is a valid assumption. I think this can indeed be the case for many types of generative processes. But the question then is then do the experiments strengthen that intuition. I feel not. It is not clear to me why e.g. EMNIST experiment there would be structural sparsity. EMNIST is also a very simple data set. I would expect the experiments to show that the learned independent components are useful in practical applications (see the brain signal experiments in Halva '21 for instance or in Khemakhem (iVAE) '20). I'm not saying specifically this type of real data experiments need to be introduced, but something to further highlight the strenght of this method would be helpful. Another example is to evaluate the method more thoroughly on some benchmarks from the disentanglement learning literature. There are also some claims in the paper that would be good to justify experimentally for instance you claim on line 311-313 that "This is particularly helpful in the context of self-supervised learning 311 (Von Kügelgen et al., 2021) or transfer learning (Kong et al., 2022), where latent representations are 312 modeled as a changing part and an invariant part." If this indeed the case, why not show that on data? Indeed 311 to 324 gives nice discussion and its a great shame this has not been shown experimentally as it would really take this paper to the next level.

Questions

I will use this space for further suggestions and questions: - Could you please introduce structural sparsity bit more clearly and intuitively -- if one has not read the original Zheng'22 paper then it is difficult to follow. - In estimation, please clarify: is it required that we know which groups of latent variables are independent, and which are potentially dependent? How is this exactly established in practice? - please explain in more detail how your algorithm, in practice, allows dimension reduction without assuming observation noise, and how jacobian can be computed for non-bijective transformation - What is the level of nonlinearity in the mixing functions? I dont believe this is mentioned anywhere, e.g. number of layers or similar. - ". Since the proposed condition is on the connective structure from sources to observed variables, i.e., the support of the Jacobian matrix of the mixing function, it does not require the mixing function to be of any specific algebraic form.". Please make this sentence bit more precise or explain better what does 'specific algebraic form' mean. Because structural sparsity does still limit the form of the function -- f can not longer be any arbitrary function. - "Most of these methods require auxiliary variables to be observable, such as class labels and domain indices (Hyvärinen and Morioka, 2016, 2017". H&M 2017, really only require the previous data so it's not really a big limitation and it's arguable whether this really constitutes of having auxiliary variables...I would consider moving that reference to the next sentence since it's a time-series model : "with the exceptions being those for time series...[move H&M'17 reference here]" - "The most obvious one arises from the fact that it may fail in a number of situations where the generating processes are heavily disentangled." Please explain in more detail why this may be? - "We first present the result on removing one of the major assumptions in ICA, i.e., the number of observed variables m must be equal to that of hidden sources n." This makes it sound like it hasn't been done previously in general, which of course it has been done many times previously in linear ICA (e.g eriksson and koivunen '03) and nonlinear ICA (e.g. khemakhem '20, halva '21 etc etc). so rather than saying removing majort assumption in ICA, make it specific to sparsity - as for the title: are you really "generalizing beyond structural sparsity"? as I feel structural sparsity is still the fundamental building block here. I would say you are generalizing structural sparsity in nonlinear ICA. - "This is similar to Independent Subspace Analysis (ISA) (Theis, 2006)" Either explain why you cite Theis, or cite an earlier ISA work (e.g. Hyvarinen & Hoyer, 2000)?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Authors should discuss limitations in more detail: - what are possible limitations of structural sparsity assumptions should be discussed more e.g. any scenarios where you expect it to be a poor assumption? - the point discussed in the "Weaknesses" section about the limitations in the identifiability theorems of 4.1 and 4.3 must be addressed and justified much more clearly or else those theorems should be removed - authors should note the heurestic approach of their estimation algorithm -- does GIN even have universal _function_ approximation capability? - related, discuss limitations of estimation algorithm in general. Is the algorithm guaranteed to find the correct sparsity for example? If not, then that should be pointed out as need for future works. - "However, our setting is more flexible in the sense that we do not assume all sources to be influenced by the auxiliary variable. Specifically, sources in $s_I$ are mutually independent as in the original ICA setting, while only sources in $s_D$ have access to the side information from the conditional independence given u,". This is true and a nice result of theorem 4.4, but there is the limitation that should be discussed, namely now you are making restricting assumption on _both_ the mixing function as well as on the auxiliary variables -- in some sense this is worst of both worlds (but still a nice theoretical result with possible practical uses).

Reviewer gBCZ8/10 · confidence 4/52023-07-10

Summary

This paper extends a recent result from Zheng et al 2022, which introduces an assumption they call “structural sparsity” to induce identifiability in nonlinear ICA without relying on a common (but arguably unrealistic) assumption that the observed variables are conditionally independent given observed auxiliary information. Whereas Zheng et al 2022 gave identifiability results only in the setting where the structural sparsity assumption holds perfectly and there are an equal number of sources and observed variables, this paper relaxes these assumptions in several interesting ways and gives identifiability or partial identifiability results in these more general settings. The first theoretical contribution shows identifiability under structural sparsity in the undercomplete setting, where there are more observed variables than sources. This lets them relax the usual assumption that the mixing function must be bijective, and instead only requires that the mixing function be injective. The second theoretical contribution relaxes the structural sparsity assumption to the setting where you have partial structural sparsity (it holds for a subset of sources) or partial independence of sources and shows partial identifiability under these assumptions. Here the partial dependence of sources does not need to be known. The third theoretical contribution assumes that the dependence between sources is known, and the fourth theoretical contribution assumes the sources with dependencies are conditionally independent given auxiliary variables (which is distinct from existing work because they don’t assume all sources are influenced by the auxiliary variable, just the dependent sources). They use an estimation method using a sparsity regularizer (that encourages a sparse estimated mixing function) with a flow-based generative model. They perform experiments on two simple visual datasets (Triangles and EMNIST) and perform ablations where they generate data that satisfy two combinations of assumptions for their theory, compared to a base setting that does not satisfy their assumptions. Following existing work, they use MCC as their metric and their models achieve higher MCC when the assumptions are satisfied.

Strengths

- Overall, this is a very interesting paper and makes novel contributions in what I think is an interesting setting: using sparsity to induce identifiability in nonlinear ICA. - They clearly motivate why relaxing each assumption makes the assumptions more realistic. - I agree that the “conditional independence given auxiliary information” assumption that is common in the literature is not a great assumption, and I’m happy to see recent work removing or reducing this assumption. - They don’t require distributional assumptions. - It is well-written and well-structured overall. It is very clear what the prior work accomplishes and what the contributions are.

Weaknesses

- There could be more experiments in realistic settings. (However, given the strength of the theoretical contributions in this paper, I think the paper should be accepted as is.) Minor comments on the writing (did not affect score): - Line 84-85: You say “part of the sources can be grouped into irreducible independent subgroups…”, but “irreducible subgroup” is a term in algebra with a specific meaning. You could avoid this “collision” by saying “irreducible independent subgroupings” or something similar. - Line 156: You start a sentence with the word “Differently, …” which sounds strange. You could say “In contrast, …” instead. - Line 167: “While this removes the restriction of bijectivity between sources and observed variables, it remains uncertain as to whether Structural Sparsity holds in general, particularly for all sources in a universal way.” - This sentence is confusing - you are saying it is uncertain whether Structural Sparsity holds in general, but Structural Sparsity is one of your assumptions. Are you saying it is uncertain whether Structural Sparsity is a reasonable assumption, based on whether it is likely to be satisfied on real-world data? - Line 177: It’s also weird to start this sentence with “Differently”. - Multiple lines: You start a handful of sentences with “Besides, …” and each time that is not really the word you mean. You should rethink how each of these sentences connects to the previous sentences and find the appropriate word for each case.

Questions

- Are you aware of Lachapelle et al 2022b? See https://arxiv.org/pdf/2207.07732.pdf. Lachapelle et al 2022a uses mechanism sparsity to induce permutation identifiability, but Lachapelle et al 2022b extends this approach to the partial identifiability setting. It would be (1) worth mentioning in the Introduction section that Lachapelle et al 2022a introduced the idea to use sparsity to induce identifiability, which inspired the approach of Zheng et al. 2022 (as stated in the text of Zheng et al 2022, see Section 3.1 of that paper), and (2) to cite Lachapelle et al 2022b as prior work using sparsity for partial identifiability (though in a distinct setting from your results as it relies on conditional independence given observed auxiliary variables).

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

- Their experiments are only on visual disentanglement tasks and there are many other interesting disentanglement or other tasks that would be interesting to see in future work. - No concerns about negative societal impacts.

Reviewer jekt7/10 · confidence 3/52023-07-27

Summary

The paper introduces more flexible ways to perform nonlinear independent component analysis (nonlinear ICA). Nonlinear ICA involves identifying the sources s from the observed x when both s and x are related by x = f(s) and f is a nonlinear function. Previous work has developed a method for this problem under a strict structural sparsity assumption that the s's and x's are one-to-one and onto, and all the s's are independent of each other. Current work provides theorems that relax the assumption in various ways including: (1) undercompletness--there can be more observed variables x than sources s; (2) partial sparsity--only a subset of all the s's may map to x's; (3) source dependence--all sources s do not have to be statistically dependent on each other; and (4) flexible grouping structures--the possibility that some of the sources can be partitioned into independent subgroups of sources. There are experiments on synthetic and real world datasets that show the effectivness of their approach.

Strengths

This paper is original because it introduces novel approaches, as far as I know, that extend the situations where nonlinear ICA can be applied. The paper exhibits good quality in various ways. First, there are various theoretical results included in the paper that extend the cases where nonlinear ICA can be applied. Second, there are also results in several experimental settings that back up the theory. The paper is mostly clear in its explanations. In terms of significance, extending the situations where nonlinear ICA can be applied is an accomplishment.

Weaknesses

While in theory extending the cases where nonlinear ICA can be applied is a strength, because there wasn't any empirical qualitative evaluation of how this approach compares to other approaches in disentangling the sources, it is not clear how significant this work is. It does not have to be a comparison of how well it disentangles sources; it could be comparing them on some other application, such as how well they extract features that are useful for classification, for example. It does not even have to be comparing this paper's approach to previous approaches; it could be comparing the different extensions of nonlinear ICA presented in this paper. Also, as pointed out by the authors, another limitation of this work is that the experimental results were only on visual datasets but not on other modalities.

Questions

While the undercompleteness result appears to me to be unique to this paper, (Zheng et al 2022) also has an undercompletness result. This paper is written so that it sounds like (Zheng et al 2022) has no undercompletness result. It would be nice if this situation could be explained or clarified. It was a bit confusing that on line 113, A is defined as a set of natural number tuples but on line 103 A is defined as a subset of natural numbers. I think on line 114 that A_{:,j} := { i | (i,j) \in S } should really be A_{:,j} := { i | (i,j) \in A }.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

Yes.

Reviewer m3QW2023-08-12

I would like to thank the authors for their clarifications.

Authorsrebuttal2023-08-16

Thank you once more for your time and suggestions. Since the discussion window narrows, might we kindly ask if our clarification has resolved the potential confusion, especially in **Q2 & A2**? Your further feedback is deeply appreciated.

Authorsrebuttal2023-08-19

Sorry for the repeated reminders. As the discussion will end in **48 hours**, would you mind kindly checking if you have any further questions? For instance, we have clarified the **implication of Thms. 4.1 and 4.3**, and you have mentioned in the second point of weakness that > "if I have understood correctly -- please correct me if I'm wrong and I'll adjust my score accordingly" As all reviewers noted, our results are important to the community in a timely manner. Thus, we would like to try our best to address any potential confusion given the opportunity for discussion.

Authorsrebuttal2023-08-19

Mistake in updating the rating?

We have been eagerly looking forward to your feedback on our detailed response. However, without seeing your feedback, we noticed that the rating by you was changed from 4 (Borderline reject) to 3 (Reject). We therefore are wondering whether it was intended to be changed this way or just a mistake. If you have any further questions or concerns, please kindly let us know. Your feedback will be appreciated.

Reviewer jekt2023-08-16

I have read your response. Thank you for preparing it. It has clarified the meaning of certain passages in the paper. Maybe it would be an even better paper if the theory could tell you whether to use a certain identifiability approach given a particular set of empirical data, rather than having to perform ablation studies, but the current paper as it is does break new ground.

Reviewer gBCZ2023-08-19

Reviewer response to rebuttal

Thanks to the authors for the thorough response to my comments and for including the additional plot in the one-page pdf. You've addressed all the questions and suggestions for improvements in my review. I will keep my score of 8.

Authorsrebuttal2023-08-19

Thanks for your reply. We didn't find any feedback from you except the score change. Would you mind kindly checking if we are the readers of that comment?

Reviewer nmMP2023-08-19

I am not sure why but it did indeed look like you were missing from 'readers'. I also couldnt edit it. So here I am re-sending it: I have now read all the reviews and rebuttals carefully. I feel like I have no choice but to reduce the rating by one point from '4' to '3' as it better reflects my current opinion of the paper. Overall the paper makes interesting theoretical contribution but the results are not as novel as the author claims, and a lot of them are to be expected as combinations of previous results (e..g sparsity and auxiliary variables combined). The experimental validation lacks breadth. Additionally the following are further problems that justify my new rating: - One of the other reviews made me realize that undercompleteness was already discussed in [1]. This result is not properly acknowledged in this paper at all. The 'introduction' is written as if the authors are presenting the combination of undercompleteness and structural sparsity for the first time. For example, in the introduction you write: *"Unfortunately, Zheng et al. (2022) require Structural Sparsity to hold for all sources in order to provide any identifiability guarantee...[paragraph change] Besides partial sparsity, identifiability with Structural Sparsity also fails with the undercompleteness (more observed variables than sources) and/or partial source dependence (potential dependence among some hidden sources)."*. This makes it look like the structural sparsity in Zheng would fail in undercomplete case and does not at all acknowledge that they in fact *do* provide result for undercomplete case. This is not good, and the introduction **must** be changed to properly attribute for this. Currently you only make a comment on Zheng's work on undercompleteness at the end of section 3 (l. 158). To raise my score, I expect to see exact changes the author will implement. - **with regards to author's A2**: You write *"In Thm. 4.1, the identifiability of $S_d$ up to an invertible transformation means that it will not be mixed with sources in $S_I$ after estimation."*. Yes I agree that your are identifiably disentangling $S_d$ and $S_I$ from each other however saying you 'identify $S_d$ up to invertible transformation' is non-sensical in that that the $\hat{S}_d$ that you estimate have no guarantee of being related in any reasonable way to the ground-truth $S_d$ (so completely unidentified). Further, you don't even say anything about $S_I$ in Thm. 4.1. But merely: "Then $s_D$ is identifiable up to an invertible transformation.". I expect to hear of the exact changes you would make to Thms 4.1 - 4.3 so that the reader gets a better picture of what these theorems *really* mean. - **wrt. A3:** You should include a short explanation of this in the paper (the part about universal approximation and GIN), if it's not already there. - **wrt. A4:** If I understand correctly, you essentially claim here that you don't need more experiments because previous literature has shown identifiability in all kinds of situations and you provide explanation of that. I think you are being far too generous to yourselves here. There are million different possible inductive biases that could potentially explain these observations (ofc. including your work too) but it is not good enough reason in my opinion for the lacking empirical side. I am happy to revise my score if these issues have good resolution. *** [1] Zheng et al. (2022), "On the Identifiability of Nonlinear ICA: Sparsity and Beyond"

Authorsrebuttal2023-08-19

Thanks so much for your further suggestions (1/2)

Thank you very much for your further comment. We respectfully believe that there exist some misunderstandings that could be addressed. We are very glad that you provide us the opportunity to further clarify and highlight them. In light of your suggestions, we have further incorporated changes in the updated version: **Q12:** The novelty compared to the result about undercompleteness in [1]. **A12:** Thanks a lot for your suggestion, which let us realize again the necessity of clarification on that earlier in the introduction. We must emphasize that **[1] did not give any identifiability result on the undercomplete nonlinear ICA at all**. As discussed in (1) A2 in the response to reviewer jekt, (2) A9 in the response to reviewer qJ69, and (3) L159-160, [1] did not give an identifiability result but only provide a way to distinguish spurious solutions due to rotated-Gaussian MPA (Thm. 3 in [1]). Thus, identifiability with Structural Sparsity and undercompleteness was not established in [1]. Since this misunderstanding can be clearly avoided by highlighting the difference between proving identifiability and distinguishing a specific spurious solution, we have added additional detailed discussions **earlier in the introduction and abstract** to avoid potential confusion, as mentioned in the responses to other reviewers. Specifically, the related sentences in the **introduction (after L81)** have been added as follows: > ”We would like to highlight that [1] provided guarantees to avoid the spurious solution of rotated-Gaussian MPA in the undercomplete case with structural sparsity. However, it is not an identifiabilty result since there exist numerous other spurious solutions in nonlinear ICA (see, e.g., the Darmois construction). Thus, the identifiability with undercompleteness and partial sparsity is still an open question, which is one of the motivations of our work.” Furthermore, we have added the following sentence in the **abstract (after L9)**: > ”Structural Sparisty has been introduced before to avoid the spurious solution of “rotated-Gaussian MPA” in the undercomplete case, but the identifiability has not been provided.” We hope the updated text could clarify the potential confusion. Thanks again for highlighting the necessity of that. If you have other related concerns, please kindly let us know. **Q13:** The exact changes on Thms 4.1 and 4.3. **A13:** We sincerely appreciate your valuable suggestions on improving the presentation and your agreement on the implications of these theorems. Regarding the presentation, we have replaced the related sentences with *“$\mathbf{s}_D$ is block-wise identifiable”*. We would like to note that the usage of the term “block-wise identifiable” originates from previous works on identifiably disentangling the changing style from images across domains (e.g., Thm. 4.2 in [2], Thm. 4.2 in [3]), which is similar to our definition of $\mathbf{s}_D$ since its distribution does change across different values of $\mathbf{u}$. We have also highlighted these works for specifying the reference of the definition. **Q14:** Include a short explanation of the universal approximation and GIN in the paper. **A14:** Thanks so much for the suggestion. It has been added, specifically after L345, as follows: > ”It is worth noting that, according to [4], coupling-based flows (e.g., GIN) are universal diffeomorphism approximators. The volume-preserving nature of GIN does not hinder it from validating our theorems, as rescaling is one of the allowed indeterminacies after identification.”

Authorsrebuttal2023-08-19

Thanks so much for your further suggestions (2/2)

**Q15:** Comments on the breath of real-world experiments. **A15:** We fully agree with you that there are *millions of different possible inductive biases* for the true hidden generating process of real-world datasets. We also agree that, besides the visual disentanglement task, there are various other exciting applications that could potentially benefit from our theory and worth exploring (e.g., the brain's signals as you mentioned). The lack of more applications is indeed a limitation of our work (as highlighted in the conclusion, L396-397), and we are very grateful for your kind suggestions on what could be the next steps. At the same time, since it is impossible to know the structure of the real-world unknown data-generating process, our theory can only be rigorously **validated** by ablation studies (via simulated data), which we have done as part of the experiments. The asymptotic guarantee has also been further validated by the experiments varying the sample size in the attached PDF in the global response. The experiments on the image datasets are indeed not very complicated, and we aim to use these results as potential, explainable, illustrations of the application scenarios, complementing our validations of the theory. Apologies if there is any potential confusion in our previous response regarding that. We do agree with you that there are more exciting real-world applications that could take this paper to the next level, and we are very grateful for your specific suggestions on them. --- Last but not least, we genuinely appreciate the time and effort you dedicated to reviewing our manuscript. We greatly value this opportunity for clarification and are thankful for the chance you provided to further improve the presentation. --- [1] Zheng et al. "On the identifiability of nonlinear ICA: sparsity and beyond." [2] Kong et al. “Partial identifiability of domain adaptation.” [3] Von Kügelgen et al. "Self-supervised learning with data augmentations provably isolates content from style." [4] Teshima et al. "Coupling-based invertible neural networks are universal diffeomorphism approximators."

Reviewer nmMP2023-08-20

**With respect to A12:** Yes that's much better and helps to avoid confusion with the Zheng paper -- it also makes your contribution and its novelty much clearer! **Wrt. A13:** Yes I am aware of the term 'block-wise identifiable' and I do think this is better but I do find it bit weird when there is no mention of $S_I$. Would you consider perhaps adding explanation under the theorem just to explain what is meant by this block-wise identifiability and explain that it does *not* mean that $S_D$ are identified element-wise? **Wrt. A14+A15:** Thanks.

Authorsrebuttal2023-08-20

We sincerely appreciate your further feedback

Thank you so much for your reply. Wrt. A13, in addition to the definition of “block-wise identifiability” before Thm. 4.1, in light of your suggestions, we have now also added the following text after L208 (directly following Thm. 4.1): > “We would like to highlight that **the block-wise identifiability of $\mathbf{s}_D$ does not mean that sources in $\mathbf{s}_D$ are element-wise identifiable** (i.e., identifiable up to an element-wise invertible transformation and a permutation). Instead, it only guarantees that sources in $\mathbf{s}_D$ will not be mixed with sources outside of $\mathbf{s}_D$ (e.g., $\mathbf{s}_I$ in Eq. 3), which might be helpful in scenarios such as disentangling changing styles across images from different domains for the purpose of finding the block of style variables, but not necessarily each individual style variable.” We hope that, by incorporating this, readers can get a clearer picture of the results. Thanks again for your constructive suggestion! We are looking forward to your kind feedback.

Authorsrebuttal2023-08-20

We are very grateful for all of your constructive and insightful suggestions! We believe that the manuscript has been improved a lot with your help.

Program Chairsdecision2023-09-21

Decision

Accept (oral)

© 2026 NYSGPT2525 LLC