We tackle the problems of latent variables identification and ``out-of-support'' image generation in representation learning. We show that both are possible for a class of decoders that we call additive, which are reminiscent of decoders used for object-centric representation learning (OCRL) and well suited for images that can be decomposed as a sum of object-specific images. We provide conditions under which exactly solving the reconstruction problem using an additive decoder is guaranteed to identify the blocks of latent variables up to permutation and block-wise invertible transformations. This guarantee relies only on very weak assumptions about the distribution of the latent factors, which might present statistical dependencies and have an almost arbitrarily shaped support. Our result provides a new setting where nonlinear independent component analysis (ICA) is possible and adds to our theoretical understanding of OCRL methods. We also show theoretically that additive decoders can generate novel images by recombining observed factors of variations in novel ways, an ability we refer to as Cartesian-product extrapolation. We show empirically that additivity is crucial for both identifiability and extrapolation on simulated data.
Paper
Similar papers
Peer review
Summary
This paper develops a theory to show how an additive decoder may be able to disentangle an image composed of several components. The proposed theory also shows how an additive decoder may produce novel images, possibly providing insights about the process performed by generative models. Besides the theory, paper also performs some experiments showing that additive decoders may have certain advantages in performing those tasks.
Strengths
Paper is well written. The topic and the experiments are interesting. It has a broad and practical view. Literature review is relatively good. The proposed theory might turn into a useful contribution for the research community.
Weaknesses
Experiments are interesting but limited and disconnected from existing experiments in the literature, in my view. ----------------- There are a few publications that are relevant but not cited: –Li, N., Raza, M.A., Hu, W., Sun, Z. and Fisher, R, Object-centric representation learning with generative spatial-temporal factorization, NeurIPS 2021. –Yoon, J., Wu, Y.F., Bae, H. and Ahn, S., An investigation into pre-training object-centric representations for reinforcement learning, ICML 2023. Both of the above papers have experiments on images that may be decomposed in the authors’ additive scenario and possibly be used as a baseline for comparison. ----------------- “Reasonableness”, mentioned under section 3.2, seems to be an unclear definition underlying a significant portion of the theory. It appears that authors attempt to define the reasonableness, yet, it is not clear what is the difference between $Z^{test}$ and $Z^{train}$. Given the bijective assumption, the difference between $Z^{test}$ and $Z^{train}$ should refer to a specific region in the domain and range of the function. Yet, it is not clear what that difference is. I understand it is hard to define the limits of the underlying manifold of relevant images - that is exactly the heart of the difficulty in developing useful theory for deep learning. Could authors expand on their “reasonable” assumption? Perhaps identifying the boundaries of the “reasonable” manifold is hard, but it may be helpful to describe and contrast what is unreasonable. Perhaps providing a discussion and citing some previous studies on the underlying manifold of images would be helpful as well, e.g.: Cohen, U., Chung, S., Lee, D.D. and Sompolinsky, H., 2020. Separability and geometry of object manifolds in deep neural networks. Nature Communications, 11(1), p.746. ----------------- There is an extrapolation study and a dataset called VAEC from the paper below. Webb, T., Dulberg, Z., Frankland, S., Petrov, A., O’Reilly, R. and Cohen, J., Learning representations that support extrapolation. ICML 2020. The images in the VAEC dataset are designed as an extrapolation task. Do authors think this task can fit into their framework? ----------------- It seems that the notation D (for the Jacobian) is only defined in the appendix under Table 2. Since the notation is used in the main body of the paper, it would be useful to define it there. If instead of D, $\nabla$ was used for the Jacobian, I would have inferred what authors mean by it. However, I was not sure about D until I found it in the appendix. ----------------- For assumption 2, it may be better to use “linearly independent” instead of “independent”. Moreover, it might be better to explain in words that: this assumption is requiring the … matrices, to have full column rank. Overall, the notion formalized in assumption 2 seems strange to me. Are authors familiar with the notion of curvature for functions and manifolds? Is there any precedence for the notion of nonlinearity defined under assumption 2 for any class of functions? I am not sure how authors’ notion of nonlinearity relates to known notions of nonlinearity/curvature, for example, the notions of curvature in differential geometry. Why should assumption 2 be satisfied over the entire manifold? In its current form, assumption 2 stacks the first and second order derivatives together and then requires that the stacked matrix to have full column rank. It is not clear to me why stacking of these matrices is necessary. If the unrolled second derivatives have full column rank, would it be necessary for the first derivatives to have full column rank? If each of the first and second order derivatives, individually have full column rank, would that be sufficient? ----------------- This is not a weakness of the paper, but perhaps worth mentioning. It is common in the literature to use x as the input to a function/model, and use y or z as the output of a function/model. However, this paper uses z to denote the inputs and x to denote the output. This may sometimes be a bit confusing. But that is the authors’ choice, and this is just feedback.
Questions
Please see the questions under weaknesses, especially the questions about assumption 2. (Assumption 2 and what is built on it is my main concern about this work. I hope authors can be more clear and more convincing about their approach.) Have authors considered expanding their own experiments to something more sophisticated? For example, the shape of objects could be not just circles, but several other types: diamonds, triangles, rectangles, etc. The size of the objects could vary as well. Do authors think their extrapolation method can be applied to the VAEC dataset (or some modification of it)? The point is to connect this paper’s experiments to existing experiments in the literature. For example, when an object in the VAEC dataset is enlarged, the enlarged object can be considered an addition of two smaller objects which would fit the authors’ additive framework. Currently, this paper’s experiments seem to be isolated from the literature. If VAEC is not suitable, authors may want to consider some other datasets from the literature.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
I did not see a discussion on limitations.
Summary
The paper analyses the statistical identifiability of latent variables in an autoencoder with a so-called additive decoder. It is shown that under this class of decoders, the blocks of latent dimensions associated with the additive decoder can be identified. This result is further related to the ability of a model to extrapolate.
Strengths
(These may be subject to change depending on the answers to my questions below) 1. The paper is generally very well written in terms of giving intuitions behind the presented math (see the counter view in Weakness 1 below). 2. The paper touches on an important topic (identifiability), and the focus on additive decoders is both interesting, relevant, and novel. 3. The paper does a nice job of providing proof sketches, which is helpful since space constraints prevent the authors from including actual proofs in the main text. 4. The paper does a very nice job of connecting assumptions and results to existing work in various branches of the literature. This is very helpful. 5. Finally, I want to emphasize that the theoretical findings are both novel and interesting.
Weaknesses
(These may be subject to change depending on the answers to my questions below) 1. The mathematics is often phrased sufficiently convoluted that the phrased intuitions are required (see Strength 1 above). E.g. I found definition 3 to be nearly unreadable, and I could not verify if the clearly phrased intuitions (lines 177-178) actually describe the math (I trust that it does, but it's a problem that it is so difficult to verify). 2. The 'additive decoder' construction seems quite similar to mixture models for which decades of work exist regarding identifiability. I was surprised to not see this link even briefly touched upon. 3. I found the extrapolation part of the paper to be less convincing than the identifiability part. Bluntly put, I got lost in the many assumptions made (Corollary 3 holds under the assumptions of Theorem 2, which hold under the assumptions of Theorem 1 which holds under Assumption 1) that I was unable to tell which were the important assumptions for the particular corollary. Thus, I lost my intuition, and the following discussion (Lines 381-304) seems rather speculative. Fere I struggle to determine what's what.
Questions
1. It seems to me that additive decoders are effectively a form of mixture model with non-trivial components. Here, we know quite a bit about conditions under which components can be identified up to permutation. This seems quite similar to the presented results. Can you elaborate on this connection? Am I misunderstanding something? 2. In definition 2, you write "let $\mathcal{B}$ be a partition..." Should this have been $\mathcal{B}$ be the set of partitions..."? Otherwise, I struggle to understand what $B \in \mathcal{B}$ actually means. But perhaps I misunderstood something. 3. Can we agree that $v := f^{-1} \circ \hat{f}$ is a diffeomorphism simply because assumption 1 states that $f$ is a diffeomorphism or is there more to this? (I ask as the statement about $v$ appears in several places in the paper) 4. How do you ensure that $f$ actually is a diffeomorphism? Here I mainly very about $f$ self-intersecting. I can see how it is easy to ensure that $f$ is an immersion, but in my reading of the paper you seem to require that $f$ is an embedding. Did I get this part right? 5. I did not understand Assumption 2 at all. Can you explain it to me? 6. In Theorem 2 it is assumed that $f$ is injective, which seems like a rather strong assumption. Can this be loosened? 7. In line 271 it is stated $Z$ is *typically* a subset of $CPE(Z)$. What is meant by "typically"? 8. I didn't quite understand the motivation behind the Cartesian-product extension (CPE). According to the intuition of Fig. 3, the CPE makes an axis-aligned extension, but isn't the entire issue regarding identifiability that we cannot assign much meaning to the axes (the axes concern the parametrization and not the underlying support)? Then, why is it natural to extend along the axes? (To be clear, I do see the point of making an extension when studying extrapolation, so my question is more why the said extension should be axis-aligned).
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
There is no need for a discussion regarding societal impact, etc., in a purely theoretical paper such as the present. I wish the paper had had a greater discussion regarding the many assumptions made throughout the paper. Such a discussion would be akin to a "limitations" section often found in more empirical/methodological papers, and I don't see why a theoretical paper should not openly be discussing the limitations of the analysis (i.e. if the assumptions are appropriate).
Summary
This paper presents the identification theory for the additive mixing function. Specifically, they transfer the existing nonlinear ICA conditions from the distribution (i.e., sufficient variability) to the nonlinear mixing function (i.e., sufficient nonlinearity). Under this model, they make the connection to extrapolation for generative models. Synthetic data experiments are designed to demonstrate their arguments.
Strengths
1. This paper is well written — both the assumptions and the implications are adequately discussed and thus easy to understand. 2. The block-wise identification result is novel as a result of the sufficient nonlinearity condition inspired by the sufficient variability condition in prior work. 3. The connection to the extrapolation is interesting and yields a valuable understanding of current large models.
Weaknesses
1. Some key assumptions, although discussed, are still evasive in their restrictiveness. The most notable is the sufficient nonlinearity assumption, which appears very restrictive. How to enforce this for the estimation model is challenging. 2. The primary assumption, namely additivity, can be very stringent. It is hard to believe this would hold for any realistic data-generating process. This also somehow trivializes the significance of the extrapolation part. I would also like to learn about the relation to recent work [1]. 3. The experimental results lack detailed explanation. I struggle to make sense of the visualizations: what are the colored shades mean in Figure 4, and what does the color mean for the dots? Are the generating processes identical in the additive and the non-additive cases, i.e., does the only difference lie in the estimation model? [1]. https://arxiv.org/abs/2305.14229
Questions
I would like to learn about the authors' response to the weaknesses listed above, which may give me a clearer perspective on the paper's contribution.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
2 fair
Limitations
Please see the weakness section.
Summary
This paper extends the recently popular approach of constraining the nonlinear function to achieve identifiable disentanglement. Here the idea is that $\mathbf{f}$ is additive i.e. made of constituent functions that operate independently on non-overlapping partitions of the latent function. Identifiable disentanglement is achieved in this situation with very mild conditions on the latent distribution. These additive decoders can be seen as rudimentary version of the decoders used in object-centric representation learning and thus help explain their generalization performance. In particular, the paper shows that by exploring the full cartesian product of the latent symbols the model can generate images that are out of the training images' support (called cartesian-product extrapolation).
Strengths
- very well written paper with clear and intuitive examples despite the technical topic - a novel way in which we can understand disentanglement by restricting the nonlinearity from the point of view of OCRL is a nice new angle to this increasingly popular field - by making assumption about the structure of $f$ the authors are able make very mild assumptions about the distribution of the latent factors which is in contrast to the much more distribution - additive decoders and the relevant results here provide a nice simple baseline model upon which future works can build more realistic ocrl models - extrapolation guarantees is a nice addition, something previous works have been missing, and is something that hopefully will be adopted by the community (though it's unclear how that could be done; see below) - assumption of sufficient nonlinearity is an insightful result and its connection to previous literature is nicely illustrated (albeit in the appendix)
Weaknesses
Since the latent variables have only very mild restrictions, the price is paid by having fairly strong restriction on the 'mixing' function i.e. block-wise additivity (likely problematic in many realistic situations such as images with occlusion) + requirements on nonlinearity. This is likely useful for OCRL (e.g. scene mixtures) but in general may be very restrictive in the more general nonlinear ICA/disentanglement and also abstracts away from the desired goal of traditional nonlinear ICA where the aim is to separate out sources from 'heavily mixed' signals. Furthermore, the additivity also leads to a slightly problematic level identifiability results in the sense that each partition-block $z_b$ is identified only up to arbitrary invertible nonlinear transformation (equation 8. & 9. and $v_B$ specifically) -- this can be interpreted as there being a nonlinear ICA / mixing problem completely *unsolved* and thus unidentified for each block. So even though we are no longer left with the generic unidentifiability problem of $x = f(z)$, we are still left with complete unidentifiability in each block $B$, $ x_B = f_B (z_B)$. It feels like the term "partition-respecting permutation" and "B-disentanglement" are thus very specific to this approach and not really corresponding to the commonly accepted definitions of disentanglement in literature. Indeed the authors write "Thus, B-disentanglement means that the blocks of latent dimensions zB are disentangled from one 177 another, but that variables within a given block might remain entangled." Related to this, for OCRL the model is quite simple, as admitted by the authors, and while they indeed may provide a good baseline (as mentioned above) it is hard to be sure whether the theorems fully explain their performance especially given these concerns about block-wise unidentifiability. It would have been nice to see more extensive experiments and especially on more complex data.
Questions
Do you believe the block-wise unidentifiability/entanglement is undesirable? Or do you believe my concerns above are not a problem? I'm talking practically speaking. What do you think is the impact of this on CPE? What if the blocks could also be fully disentangled -- what would this change (e.g. in CPE?). Again I'm interested in your thoughts on practical implications rather than theory. You say that: "is “imitating” a block-specific ground-truth decoder". Could you please explain what you exactly mean by imitation? I find that quite vague. You write that "We believe it illustrates the fact that disentanglement alone is not sufficient to enable extrapolation and that one needs to restrict the hypothesis class of decoders in some way." These sound like very generic statements -- do you have confidence that they hold much beyond this setting? "our in this work to understand “out-of-support” generation is a step towards understanding theoretically why modern generative models such as DALLE-2 [42] and 53 Stable Diffusion [43] can be creative" Could you please explain this more specifically? Do you believe additiveness to be important to this?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
4 excellent
Contribution
3 good
Limitations
There is some good discussion of limitations in the paper e.g. "additive decoders make intuitive sense for OCRL, they are not expressive enough to represent the “masked decoders” typically used in", "Additionally, this parameter sharing across f (B) enables modern methods to have a variable number of objects across samples, an important practical point our theory does not cover." Appendix also has nice examples of what happens when some assumptions are violated.
Summary
Motivated by the problem of disentanglement in generative models, this paper proposes a novel decoder architecture, so-called additive decoders, based on the addition of block-wise latent variables, where blocks of latents correspond to semantic factors. An extrapolation property of their decoder is demonstrated, which essentially consists of forming new products of latent blocks unseen at train time. A theoretical analysis of their decoder is carried out. Several new definitions are made, as well as few necessary assumptions. Two theorems are presented on "local disentanglement" and "global disentanglement". An empirical investigation based on synthetic image data, consisting of two balls randomly placed, is carried out, showing a case where their decoder improves upon a baseline case.
Strengths
The paper addresses the problem of disentanglement in an interesting way: through blocks of latent variables which contribute additively to the overall decoding. This is a natural and reasonable idea. The motivation for this paper is good: restrict the decoder to be additive to address the problem of latent variable identifiability. The proof technique appears to build on the well-established results (Hyvärinen, AISTATS 2019).
Weaknesses
The paper is quite technical. This makes it challenging to ensure all technical details are correct. I did not find any errors. The authors write that this work may help explain the "creativity" of mainstream generative models like DALLE-2 and Stale Diffusion. It is not so clear to me that their analysis will be helpful. Regarding Assumption 2, which the authors write is "key" for Theorem 2, I am not clear on how realistic this assumption is in practice. I understand that it is motivated by similar assumptions made in the ICA literature. However I come away with the feeling that it may not have much practical relevance. The authors give a toy numerical example in Example 3, which is helpful. But for instance, has this assumption been verified for the additive decoders used in the Experiments (Section 4)? The proposed additive decoder cannot handle occlusions, as noted by the authors in Section 4, and discussed in the appendix. This is a limitation, since real images may have occlusions. Only small-scale synthetic data are included. It doesn't appear to me that training on larger scale data is feasible, but more discussion of this would be helpful. It's reasonable to develop theory on simple cases, perhaps even necessary in this cases, but we should be clear whether scaling the results can be expected.
Questions
Can the authors further justify their statement that this work "has the potential of expanding our creativity in generative model" (line 368). Can assumption 2 be verified for the decoders in used in the Experiments (Section 4)? Is training on larger-scale data possible with this method? If not, why not?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Unfortunately I did not find any discussion of limitations in the main text. I don't think this paper requires a discussion on societal impact.
I believe the authors' response and the discussions included in the appendix address most of my questions/concerns. The 1-page pdf file also provides useful clarifications about the authors’ work. The remaining point, in my view, is the requirement in assumption 2 which is not related to known notions of nonlinearity. However, authors explain in the rebuttal (and in the appendix) that this assumption somehow relates to previous assumptions in the OCRL literature. They further explain in the pdf that this assumption is satisfied for the models they have used in their experiments. As this is a conference paper, I find these arguments convincing and I’m raising my score. I suggest authors expand further on their examples 2 and 3, and explain the implications of their assumption 2 for neural networks, e.g., which models would not satisfy this requirement, which models would satisfy it, what is the minimal neural network architecture that would satisfy the requirement in assumption 2, etc.
Thank you for engaging with the rebuttal and adjusting your evaluation. With the additional space, we should be able to add details about example 2 and 3, thanks for the suggestion. Clarification: Assumption 2 is not about the model used for learning. It is an assumption about the data-generating process. Here, since the dataset is synthetic, we can test the assumption, however in general it might not be possible.
Thank you for the follow-up
Thank you for the follow. Indeed, your analysis was spot on for my confusion regarding links to mixture models. I will increase my score.
I found the author's rebuttal convincing. The new numerical experiment on explicitly verifying Assumption 2 is helpful. At least for toy data with well-conditioned matrices, this assumption can be verified in practice. I will raise my score.
I appreciate the detailed response from the authors -- many thanks! I have raised my rating to reflect this. I would be interested in learning from the authors about the relationship between additive and compositional functions, if they happen to have further updates on this.
Thanks for engaging with our rebuttal! Regarding the relationship between additive and compositional decoders, we now almost have a complete proof that compositional implies additive. We are stuck on a small technical detail. In the worst case, adding a mild regularity assumption should do the trick. Essentially, we (almost) showed that a compositional decoder has a Hessian with a single nonzero element that lies on its diagonal. This of course means that they have diagonal Hessians and thus are additive (Appendix A.2 shows that additivity is equivalent to diagonal Hessian). This is interesting has it makes the connection between both function classes very transparent. We will make sure to give more updates before the end of the discussion period to confirm (or infirm) everything.
We now have a complete proof that **($C^2$) compositional decoders are additive**. This implies that the class of additive decoders is *strictly* more expressive than the class of $C^2$ compositional decoder ([Brady et al., 2023] assumes only $C^1$, see below for more on this). We provide a partial proof here. We’ll be glad to provide further details on request. We give a definition of compositional decoder adapted from Brady et al.. **Def:** Given a partition $\mathcal{B}$, a function $f$ is compositional w.r.t. $\mathcal{B}$ when, for all $i \in [d_z], z \in \mathbb{R}^{d_z}, B \in \mathcal{B}$, we have that $D_B f_i(z) \not= 0 \implies D_{B^c} f_i(z) = 0$, where $B^c = [d_z] \setminus B$. **Proof sketch:** Our strategy is to show that the Hessian of $f_i$ is block diagonal everywhere on $\mathbb{R}^{d_z}$ and then use Proposition 5 from Appendix A.2 to conclude that $f_i$ must be additive. For each $z_0 \in \mathbb{R}^{d_z}$, we know there exists a $B \in \mathcal{B}$ such that $D_{B^c}f_i(z_0) = 0$. In the case where $D_Bf_i(z_0) \not= 0$, we have by continuity of $Df_i$ that there exists an open neighborhood of $z_0$ on which $D_Bf_i(z) \not= 0$. By compositionality, we must also have that $D_{B^c}f_i(z) = 0$ on that neighborhood. This means that the derivative of $D_{B^c}f_i$ at $z_0$ is zero, i.e. $DD_{B^c}f_i(z_0) = 0$. Since $f$ is $C^2$, its Hessian is symmetric and thus $D_{B^c}Df_i(z_0) = 0$. We can thus conclude that $D^2f_i(z_0)$ is filled with zeros except possibly at the entries $B\times B$. Hence it is block diagonal. In the case where $D_Bf_i(z_0) = 0$, the argument is slightly more involved because we cannot necessarily take derivatives because $(D_Bf_i)^{-1}(\\{0\\})$ is a closed set (by continuity of $D_Bf_i$) and thus $z_0$ might be on its boundary. If $z_0$ is in the interior of $(D_Bf_i)^{-1}(\\{0\\})$, we can use an argument similar to above to show that $D^2 f_i(z_0) = 0$. If $z_0$ is on the boundary, by using the continuity of $D^2f_i$, we can show the Hessian is also going to be block-diagonal. (We made this last part of the argument more precise in our revision. We can provide these additional details on request.) $\blacksquare$ **We would like to reiterate the differences between both works:** - We consider additive decoders which, as we just showed, are strictly more expressive than $C^2$ compositional decoders introduced in [Brady et al., 2023]. - We assume the decoder is $C^2$ whereas [Brady et al., 2023] assumes only C^1 (which is weaker). - We consider very general domains for the latent vector z (Brady et al. study only fully supported latent vectors) - We prove additive decoders can extrapolate (their work has no discussion of extrapolation). Note that we cannot say that our identifiability result is stronger than theirs because we assume $C^2$ decoders while they assume $C^1$. This also makes the comparison between our “sufficient nonlinear” assumption (which refers to second derivatives) and their “irreducibility assumption” (which we believe to be somewhat analogous) difficult. We'd like to correct a mistake we made in a previous message: We said that "compositional" means that the Jacobian cannot have more than one nonzero entry per row. This is true only in the special case where $\mathcal{B} := \{\{1\}, ..., \{d_z\}\}, but the Brady et al. allowed for more general partitions. The definition we gave above is correct. Feel free to ask if you have any questions, we'll be happy to clarify.
Thanks for the comments -- I am mostly happen with them except below: **We believe that extending existing proof techniques to block identifiability is actually a strength, since it is more general and it includes the trivial partition {{1}, {2}, ..., {d_z}}, which would yield full disentanglement.** I am not convinced by this argument. If you have have partition {{1}, {2}, ..., {d_z}} then each $f^{(b)}(z_b)$ only takes as an input a single random variable, no? so there is complete lack of nonlinear mixing (only a linear mixture is disentangled to find the individual components). As soon as the partitions are larger than size 1, the block-specific nonlinear mixture is immediately unidentifiable. Therefore, you are not solving nonlinear ICA, you are only disentangling the partitions from each other -- there representations you learn for each block may be completely arbitrary (bijective) transformations of the ground-truth representations. Is this correct? *** Figure 4 -- as someone who has red-green color blindness, the red square of extrapolations in Figure 4 is almost invisible (i only saw it once zooming in full screen). The colors in general are poorly chosen -- there exist several colour palettes that take these issues into account. *** p.s. dodgy grammar in l.286 "i.e. it is the observation one would have obtain by evaluating"
Thanks for seriously engaging with our rebuttal, we appreciate it. Below we address your concerns. We apologize for the lengthy answer, but we felt some point required careful explanations. **"If you have have partition {{1}, {2}, ..., {d_z}} then each $f^{(b)}(z_b)$ only takes as an input a single random variable, no?"** That is correct. **"so there is complete lack of nonlinear mixing (only a linear mixture is disentangled to find the individual components)"** It depends on what is meant by "complete lack of nonlinear mixing". The resulting data-manifold can still be highly nonlinear. However, it is true that the nature of the mixing between components is limited by the additivity. I believe the point you are making here reduces to the point you initially made that "additivity is restrictive". We agree with this. Many works have considered restricted function classes to improve identifiability like [Brady et al., 2023], [Buchholz et al., 2022], [Gresele et al., 2021] and [Taleb & Jutten, 1999]. It is clear that these works as well as ours do not form a complete solution to the nonlinear ICA problem, which, in his original formulation, is known to be unsolvable [Hyvärinen & Pajunen, 1999]. That being said, some function classes will be useful for some applications, but not all. In this work we argued that additivity makes intuitive sense for object-centric representation learning (with caveats). **"As soon as the partitions are larger than size 1, the block-specific nonlinear mixture is immediately unidentifiable."** That is correct. The blocks of the partition will be disentangled from one another, but the variables within a given block can remain entangled. **"Therefore, you are not solving nonlinear ICA, you are only disentangling the partitions from each other"** The term "nonlinear ICA" sometimes mean different things in different context. Some people use it to refer to the original problem where the decoder is a general invertible map and the latents are independent [Taleb & Jutten, 1999] (which is unidentifiable). Some will use it to mean any setting where the mixing function is a general invertible map but will allow for richer latent distribution like with auxiliary variables [Khemakhem et al., 2020] or temporal dependencies [Lachapelle et al., 2022] for example. In this work, we use the term "nonlinear ICA" to mean any problem where the goal is to recover latent variables from some nonlinear mixture (which might be restricted). Of course, under this definition, additive decoders count as "nonlinear ICA" (since additive functions are nonlinear in general). Note that, although our identifiability result restricts the mixing function, we allow for much more general distribution over the latents (dependencies + general support shape) than the strictest interpretation of "nonlinear ICA" which assumes independent latents. **"there representations you learn for each block may be completely arbitrary (bijective) transformations of the ground-truth representations. Is this correct?"** Indeed, the block $\hat{z}\_B$ of a learned representation can be an arbitrary nonlinear transformation of some block of the ground-truth $z_{B'}$. **Poor choice of color in Figure 4** Sincerely sorry for the inconvenience, we'll make sure to fix this in the camera-ready version. You said that in general the colors are poorly chosen, are there any other specific places that caused trouble? [Brady et al., 2023] https://arxiv.org/abs/2305.14229 [Buchholz et al., 2022] https://arxiv.org/abs/2208.06406 [Gresele et al., 2021] https://arxiv.org/abs/2106.05200 [Taleb & Jutten, 1999] https://ieeexplore.ieee.org/document/790661 [Hyvärinen & Pajunen, 1999] https://www.cs.helsinki.fi/u/ahyvarin/papers/NN99.pdf [Khemakhem et al., 2020] https://arxiv.org/abs/1907.04809 [Lachapelle et al., 2022] https://arxiv.org/abs/2107.10098
I still find this problematic. With regards to above discussion about nonlinear ICA, and lack of identifiability of nonlinear mixtures, I can see where you are coming from, but still hold my opinion that this paper is not really doing nonlinear ICA. You say that: *"The term "nonlinear ICA" sometimes mean different things in different context. [...] Some will use it to mean any setting where the mixing function is a general invertible map but will allow for richer latent distribution like with auxiliary variables [Khemakhem et al., 2020] or temporal dependencies [Lachapelle et al., 2022] for example"* Perhaps, but using the term ICA to refer to situations where the latent variables are not independent, logically does not really make sense and is poor usage, in my opinion, given what the abbreviation stands for. In Khemakhem (iVAE paper) the latent components are conditionally independent so I feel it's justified there. As a case in point, Khemakhem et al. also have works where variables are *not* independent -- in his phd thesis he calls those works 'identifiable representation learning', which I find much more appropriate. Similarly the prefix 'nonlinear' is also not really justified when one is **not** able to recover nonlinearly mixed components, but can merely recover the partitions (symbols/objects). **To accept this paper, I would require this to be discussed more openly** Currently you write: l.177 "Thus, B-disentanglement means that the blocks of latent dimensions zB are disentangled from one another, but that variables within a given block might remain entangled." I don't think this is enough because a careless / casual reader might not easily realize how different this is from the typical nonlinear ICA. I think it would suffice however if you had a sentence or two along the lines "note that the B-disentanglement result is different from those in nonlinear ICA in that we do not recover each latent component, but rather disentangle the partitions, that is, variables within a given block might remain entangled". Key is to stress the difference to nonlinear ICA which is missing now. *Sincerely sorry for the inconvenience, we'll make sure to fix this in the camera-ready version. You said that in general the colors are poorly chosen, are there any other specific places that caused trouble?* I just meant Figure 4 in general -- it is also hard for me to tell apart the two balls as they are red and greenish or something like that. Figure 5 is fine. I appreciate that you can not see the Figures through other peoples eyes so it can always be hard to find colors that suit everyone!
Thanks for engaging with us and insisting on that point! After some thinking and discussion, we’ve come to agree with you that our usage of the term “ICA'' was too broad and that “ICA” should be reserved for methods where some form of statistical independence between the components is assumed for identification. Of course, under this definition, our approach does not qualify as ICA. To address this issue, we propose the following modifications: 1- Regarding your request to contrast more transparently with nonlinear ICA, we suggest adding this paragraph at L87 in the “Background & literature review”: **Relation to nonlinear ICA.** [Hyvärinen & Pajunen, 1999] showed that the standard nonlinear ICA problem where the observation $x$ is given by a general nonlinear transformation of *statistically independent latent factors* $z_i$ is unidentifiable. This motivated various extensions of nonlinear ICA where more structure on the factors is assumed [CITE nonlinear ICA works]. Our approach departs from the standard nonlinear ICA problem along three axes: (i) we restrict the mixing function to be additive, (ii) the factors do not have to be necessarily independent, and (iii) we can identify only the blocks $z_B$ as opposed to each $z_i$ individually up to element-wise transformations, unless $\mathcal{B} = \\{\\{1\\}, ...,\\{d_z\\}\\}$ (see Section 3.1). 2- We will also add the following clarification right after L178 in Section 3.1: “Note that, unless the partition is $\mathcal{B} = \\{\\{1\\}, …, \\{d_z\\}\\}$, this corresponds to a weaker form of disentanglement than what is typically seeked in nonlinear ICA, i.e. recovering each variable individually.” 3- We suggest replacing the following problematic sentence from the abstract: L10: “Our result provides a new setting where nonlinear independent component analysis (ICA) is possible and adds to our theoretical understanding of OCRL methods.” by “Our result adds to our theoretical understanding of OCRL methods and provides a new variation of nonlinear independent component analysis (ICA) where latent factors can be identified.” (The rationale behind keeping the keyword “ICA” in the abstract is that we believe this result will be of interest to this community.) Please let us know whether you find these modifications to be satisfactory or not. Note: We would like to clarify your point that “Similarly the prefix 'nonlinear' is also not really justified when one is not able to recover nonlinearly mixed components, but can merely recover the partitions (symbols/objects)”. Strictly speaking, a function of the form $\sum_B f^{(B)}(z_B)$ can be nonlinear in $z$, even when the partition is trivial. That being said, we agree that a case can be made that the *mixing* itself is linear in the following sense: Additive decoders with the trivial partition can be written as $f(z) = S(F(z))$ where $F(z) = [f^{(1)}(z_1), …, f^{(d_z)}(z_{d_z})] \in \mathbb{R}^{d_x \times d_z}$ and $S: \mathbb{R}^{d_x \times d_z} \rightarrow \mathbb{R}^{d_x}$ is the linear operator consisting of summing the columns of $F(z)$. Of course, $F(z)$ can be nonlinear, but it does not mix the latent factors. The mixing occurs only in $S$, which is a linear operator. So although the decoder function $f(z) = S(F(z))$ is indeed nonlinear, the *mixing step* is linear.
Decision
Accept (oral)