Summary
This work introduces a novel identifiability analysis for a hierarchical latent models where latent variables are discrete and observations are continuous. The novelty of the result resides in the fact that previous results consider mainly continuous latent variables or make stronger assumptions on the form of the latent graph. An algorithm based on the theory is proposed and tested on synthetic data. Analogies between the approach and diffusion models are drawn.
**Review summary**
Overall I believe the theoretical contribution is interesting, important and novel, but the presentation requires some non-trivial restructuring since at the moment, a significant portion of the content of the paper is relayed to the appendix which makes it very hard to follow. I also thought the section on Diffusion Models was unconvincing and a bit disconnected from the rest of the contributions. I provided some suggestions, including submitting to a venue that allows for more space, like JMLR for instance. Given this, I can only recommend borderline acceptance.
Strengths
- I believe the problem of identifiability in hierarchical latent variable models is interesting and important.
- The theory presented seems non-trivial and valuable (I did not read the appendix)
- Most identifiability results assumes continuous latent variables, so I was pleased to see further progress made in the case of discrete latents, which is much less common in the literature.
- I appreciated the high-level explanation of the proof technique between lines 223-233 which makes the connection to prior work transparent.
- The work is transparent about its limitations.
- Many examples are presented, which is helpful to understand the complex notions.
Weaknesses
**Writing**
I thought the writing was quite good and easily understandable up until Section 3.3, where quality started degrading in my opinion. It really looks like the authors were running out of space and decided to relay *a very large* portion of the content to the appendix. Here's a (probably non-exhaustive) list of important concept and contributions which were relayed to the appendix:
- t-separation
- non-negative rank
- the minimal-graph operator
- the skeleton operator
- Condition A3.15
- Algorithm 1 (this is the main practical contribution!)
- adaptive sparsity selection mechanism for capturing concepts at different levels (another practical contribution)
- The literature review.
I can understand when a proof or even when a few very technical assumptions are kept in the appendix, as long as it does not interfere with understanding what is said in the main paper. But here, all these notions are referred to in definitions and assumptions and this really makes some sections unreadable. Also, some of these notions are not standard at all, like t-separation (I'm familiar with d-seperation) or non-negative rank (I'm familiar with the standard notion of rank) and would benefit from explanations in the main text.
In addition, Algorithm 1, which is the main practical contribution, is described only in the appendix. Same thing for the adaptive sparsity selection mechanism for capturing concepts at different levels in diffusion models from Section 6.2. The literature review is in the appendix.
**Diffusion models experiments**
I appreciate the effort to include more realistic experiments in a theoretical paper, but here I felt like Sections 5 & 6 on diffusion models were disconnected from the rest of the paper… My understanding is that the authors do not apply Algorithm 1 developed so far to the diffusion model. It seems the point of these sections is to draw what I believe to be very vague connections between the assumptions of their hierarchical model and the hierarchical nature of diffusion models. Section 6 only shows that different noise level of the latent space of a diffusion model correspond to our intuitive sense of “abstract levels”. But AFAIK this is a well known observation, no? Section 6.2 introduces another algorithm with only very high-level explanations with details in appendix.
**Suggestions for improvements**
I believe this manuscript would be more suited for a journal like JMLR than for a conference. The additional space would allow the authors to present all definitions in the main text and give intuitions for their meaning (for instance, the definition of atomic cover is very dense and could benefit from more explanations and intuitions. The recursive nature of the definition makes it quite challenging to grasp IMO). This also avoid the endless back and forth between main text and appendix.
Another possibility would be to remove the section on latent diffusion models, but even then this might not be enough.
**Relatively minor points:**
- Line 130-132: I believe the estimators d_hat, c_hat, g_hat and \Gamma_hat should be defined more explicitly, given how crucial they are to the results. In this phrasing, it is not clear whether these are estimated on a finite dataset or the full population.
- Table 1 and 2 are not referred to in the main text.
- Condition 3.1: The notion of splitting a latent variable is not properly explained.
- Condition 3.3: By definition, the support of a random variable is closed. See for example: https://en.wikipedia.org/wiki/Support_(mathematics)#In_probability_and_measure_theory . The only subsets of Rn that are both open and closed are the empty set and Rn itself. I’m guessing the authors were hoping to include more sets in their theory. It might be possible by assuming the set is “regular closed”, which means it is equal to the closure of its interior. This was done in a similar setting in [65].
Interesting to see that (iii) resembles the notion of G-preservation from [a] (see Definitions 11-12 and Proposition 3)
- Be careful with phrasing like line 237 “We define t-separation in Definition A3.2” as it sounds a bit like the authors are introducing this notion, but it’s not the case (source is cited properly in appendix).
- Confusion around t-separation: In Theorem 3.5, it is written “L t-separates A and B in G”, but the definition of t-sep refers to a tuple, i.e. “(L_1, L_2) t-separates A and B in G”. Not sure what the statement means.
- Text is too small in Figure 3
- Line 119: The definition of pure child was a bit confusing. In particular I thought B could contain more nodes than just the parents of A. Why not just repeat Definition A3.8 in the main text? (ne need to have a definition environment)
- Typo on line 141, V_1 or v_1 ?
Questions
Condition 3.1
- The full support condition seems a bit strong, can the author discuss what it would mean for the running example with the dog?
- The function ne(v) was not defined, this is neighbors of v, right? It’s the union of Parents and children of v, correct?
The sparsity condition of Condition 3.3(iii) seems to be crucial for disentanglement in Theorem 3.4. How does this assumption compare to other works using sparsity of the decoder for disentanglement, such as [b,c,d]? I really believe there should be a discussion comparing the graphical assumptions with those of [d].
Definition 3.6: At line 256, what is the support of a set of atomic covers? (Supp(C) ?)
**References**
[65] Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, and Simon Lacoste-Julien. Additive decoders for latent variables identification and cartesian-product extrapolation. Advances in Neural Information Processing Systems, 36, 2024.
[a] S. Lachapelle, P. R. Lopez, Y. Sharma, K. Everett, R. L. Priol, A. Lacoste, and S. Lacoste-Julien. Nonparametric partial disentanglement via mechanism sparsity: Sparse actions, interventions and sparse temporal dependencies, 2024.
[b] J. Brady, R. S. Zimmermann, Y. Sharma, B. Scholkopf, J. von Kugelgen, and W. Brendel. Provably ¨ learning object-centric representations. In Proceedings of the 40th International Conference on Machine Learning, 2023.
[c] G. Elyse Moran, D. Sridhar, Y. Wang, and D. Blei. Identifiable deep generative models via sparse decoding. Transactions on Machine Learning Research, 2022.
[d] Y. Zheng, I. Ng, and K. Zhang. On the identifiability of nonlinear ICA: Sparsity and beyond. In Advances in Neural Information Processing Systems, 2022
Limitations
Limitations were discussed properly throughout the paper.