Learning Discrete Concepts in Latent Hierarchical Models

Learning concepts from natural high-dimensional data (e.g., images) holds potential in building human-aligned and interpretable machine learning models. Despite its encouraging prospect, formalization and theoretical insights into this crucial task are still lacking. In this work, we formalize concepts as discrete latent causal variables that are related via a hierarchical causal model that encodes different abstraction levels of concepts embedded in high-dimensional data (e.g., a dog breed and its eye shapes in natural images). We formulate conditions to facilitate the identification of the proposed causal model, which reveals when learning such concepts from unsupervised data is possible. Our conditions permit complex causal hierarchical structures beyond latent trees and multi-level directed acyclic graphs in prior work and can handle high-dimensional, continuous observed variables, which is well-suited for unstructured data modalities such as images. We substantiate our theoretical claims with synthetic data experiments. Further, we discuss our theory's implications for understanding the underlying mechanisms of latent diffusion models and provide corresponding empirical evidence for our theoretical insights.

Paper

References (87)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer BR1s5/10 · confidence 2/52024-07-02

Summary

This paper studies the framework that identifies the discrete hierarchical latent variables for learning concepts from observed data examples. The proposed theory can be used to interpret the generating process of latent diffusion probabilistic models from the perspective of constructing object concepts.

Strengths

1. The paper is, in general, well-written and well-motivated and focuses on the difficult problem of capturing concepts in vision problems. 2. This work includes thorough theoretical derivation and details. 3. The illustrations of synthetic data demonstrate the applicability of the proposed method, and the results look interesting.

Weaknesses

This paper is nice to read, while I have only limited experience in such a causal hierarchical modelling area. My questions can be found below. 1. From Sec. 2, given the description that the continuous latent variables c seem to control a lower level of features of data, while in Figure 1. a, it seems to be independent to concept factor d at the same level of the hierarchical structure. Could you elaborate on the relation and difference between c and d? 2. For LD experiments, at the early steps of the diffusion process (i.e., T), Figure 5 presents controlling of bread and species (high-level) features, while the low-level (e.g., object angle, background) can remain the same. But, in Figure A8, such capability does not hold, especially for the shoes. Can the author explain the reason behind this?

Questions

Please see the weakness.

Rating

5

Confidence

2

Soundness

3

Presentation

2

Contribution

2

Limitations

Yes.

Reviewer ae895/10 · confidence 4/52024-07-10

Summary

This work introduces a novel identifiability analysis for a hierarchical latent models where latent variables are discrete and observations are continuous. The novelty of the result resides in the fact that previous results consider mainly continuous latent variables or make stronger assumptions on the form of the latent graph. An algorithm based on the theory is proposed and tested on synthetic data. Analogies between the approach and diffusion models are drawn. **Review summary** Overall I believe the theoretical contribution is interesting, important and novel, but the presentation requires some non-trivial restructuring since at the moment, a significant portion of the content of the paper is relayed to the appendix which makes it very hard to follow. I also thought the section on Diffusion Models was unconvincing and a bit disconnected from the rest of the contributions. I provided some suggestions, including submitting to a venue that allows for more space, like JMLR for instance. Given this, I can only recommend borderline acceptance.

Strengths

- I believe the problem of identifiability in hierarchical latent variable models is interesting and important. - The theory presented seems non-trivial and valuable (I did not read the appendix) - Most identifiability results assumes continuous latent variables, so I was pleased to see further progress made in the case of discrete latents, which is much less common in the literature. - I appreciated the high-level explanation of the proof technique between lines 223-233 which makes the connection to prior work transparent. - The work is transparent about its limitations. - Many examples are presented, which is helpful to understand the complex notions.

Weaknesses

**Writing** I thought the writing was quite good and easily understandable up until Section 3.3, where quality started degrading in my opinion. It really looks like the authors were running out of space and decided to relay *a very large* portion of the content to the appendix. Here's a (probably non-exhaustive) list of important concept and contributions which were relayed to the appendix: - t-separation - non-negative rank - the minimal-graph operator - the skeleton operator - Condition A3.15 - Algorithm 1 (this is the main practical contribution!) - adaptive sparsity selection mechanism for capturing concepts at different levels (another practical contribution) - The literature review. I can understand when a proof or even when a few very technical assumptions are kept in the appendix, as long as it does not interfere with understanding what is said in the main paper. But here, all these notions are referred to in definitions and assumptions and this really makes some sections unreadable. Also, some of these notions are not standard at all, like t-separation (I'm familiar with d-seperation) or non-negative rank (I'm familiar with the standard notion of rank) and would benefit from explanations in the main text. In addition, Algorithm 1, which is the main practical contribution, is described only in the appendix. Same thing for the adaptive sparsity selection mechanism for capturing concepts at different levels in diffusion models from Section 6.2. The literature review is in the appendix. **Diffusion models experiments** I appreciate the effort to include more realistic experiments in a theoretical paper, but here I felt like Sections 5 & 6 on diffusion models were disconnected from the rest of the paper… My understanding is that the authors do not apply Algorithm 1 developed so far to the diffusion model. It seems the point of these sections is to draw what I believe to be very vague connections between the assumptions of their hierarchical model and the hierarchical nature of diffusion models. Section 6 only shows that different noise level of the latent space of a diffusion model correspond to our intuitive sense of “abstract levels”. But AFAIK this is a well known observation, no? Section 6.2 introduces another algorithm with only very high-level explanations with details in appendix. **Suggestions for improvements** I believe this manuscript would be more suited for a journal like JMLR than for a conference. The additional space would allow the authors to present all definitions in the main text and give intuitions for their meaning (for instance, the definition of atomic cover is very dense and could benefit from more explanations and intuitions. The recursive nature of the definition makes it quite challenging to grasp IMO). This also avoid the endless back and forth between main text and appendix. Another possibility would be to remove the section on latent diffusion models, but even then this might not be enough. **Relatively minor points:** - Line 130-132: I believe the estimators d_hat, c_hat, g_hat and \Gamma_hat should be defined more explicitly, given how crucial they are to the results. In this phrasing, it is not clear whether these are estimated on a finite dataset or the full population. - Table 1 and 2 are not referred to in the main text. - Condition 3.1: The notion of splitting a latent variable is not properly explained. - Condition 3.3: By definition, the support of a random variable is closed. See for example: https://en.wikipedia.org/wiki/Support_(mathematics)#In_probability_and_measure_theory . The only subsets of Rn that are both open and closed are the empty set and Rn itself. I’m guessing the authors were hoping to include more sets in their theory. It might be possible by assuming the set is “regular closed”, which means it is equal to the closure of its interior. This was done in a similar setting in [65]. Interesting to see that (iii) resembles the notion of G-preservation from [a] (see Definitions 11-12 and Proposition 3) - Be careful with phrasing like line 237 “We define t-separation in Definition A3.2” as it sounds a bit like the authors are introducing this notion, but it’s not the case (source is cited properly in appendix). - Confusion around t-separation: In Theorem 3.5, it is written “L t-separates A and B in G”, but the definition of t-sep refers to a tuple, i.e. “(L_1, L_2) t-separates A and B in G”. Not sure what the statement means. - Text is too small in Figure 3 - Line 119: The definition of pure child was a bit confusing. In particular I thought B could contain more nodes than just the parents of A. Why not just repeat Definition A3.8 in the main text? (ne need to have a definition environment) - Typo on line 141, V_1 or v_1 ?

Questions

Condition 3.1 - The full support condition seems a bit strong, can the author discuss what it would mean for the running example with the dog? - The function ne(v) was not defined, this is neighbors of v, right? It’s the union of Parents and children of v, correct? The sparsity condition of Condition 3.3(iii) seems to be crucial for disentanglement in Theorem 3.4. How does this assumption compare to other works using sparsity of the decoder for disentanglement, such as [b,c,d]? I really believe there should be a discussion comparing the graphical assumptions with those of [d]. Definition 3.6: At line 256, what is the support of a set of atomic covers? (Supp(C) ?) **References** [65] Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, and Simon Lacoste-Julien. Additive decoders for latent variables identification and cartesian-product extrapolation. Advances in Neural Information Processing Systems, 36, 2024. [a] S. Lachapelle, P. R. Lopez, Y. Sharma, K. Everett, R. L. Priol, A. Lacoste, and S. Lacoste-Julien. Nonparametric partial disentanglement via mechanism sparsity: Sparse actions, interventions and sparse temporal dependencies, 2024. [b] J. Brady, R. S. Zimmermann, Y. Sharma, B. Scholkopf, J. von Kugelgen, and W. Brendel. Provably ¨ learning object-centric representations. In Proceedings of the 40th International Conference on Machine Learning, 2023. [c] G. Elyse Moran, D. Sridhar, Y. Wang, and D. Blei. Identifiable deep generative models via sparse decoding. Transactions on Machine Learning Research, 2022. [d] Y. Zheng, I. Ng, and K. Zhang. On the identifiability of nonlinear ICA: Sparsity and beyond. In Advances in Neural Information Processing Systems, 2022

Rating

5

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

Limitations were discussed properly throughout the paper.

Authorsrebuttal2024-08-14

Thank you so much for your feedback. We're really glad to hear that you think our changes can improve the manuscript by quite a bit. Your suggestions have been incredibly helpful -- thank you again!

Reviewer 7auw6/10 · confidence 3/52024-07-13

Summary

This paper introduces a theoretical framework for learning discrete concepts from high-dimensional data using latent hierarchical causal models. The key contributions are: 1) Formalizing concept learning as identifying discrete latent variables and their hierarchical causal structure from continuous observed data. 2) Providing identifiability conditions and proofs for recovering discrete latent variables and their hierarchical relationships. 3) Interpreting latent diffusion models through this hierarchical concept learning lens, with supporting empirical evidence. The work bridges theoretical causal discovery with practical deep generative models, offering new perspectives on how concepts might be learned and represented.

Strengths

1. Novel formalization of concept learning as a causal discovery problem, providing theoretical grounding for an important area of machine learning 2. Rigorous proofs for identifiability of discrete latent variables and hierarchical structures under relatively mild conditions 3. Flexible graphical conditions that allow for more complex hierarchical structures than previous work 4. Interesting connection drawn between the theoretical framework and latent diffusion models, with empirical support 5. Clear potential impact on understanding and improving deep generative models for concept learning

Weaknesses

1. The identification conditions (Condition 3.3 and 3.7) may be too restrictive for real-world scenarios. For instance, the invertibility requirement on the generating function g (Condition 3.3-ii) could be difficult to guarantee in practice, especially for complex high-dimensional data. 2. The method relies heavily on rank tests of probability tables (Theorem 3.5), which can be computationally expensive and numerically unstable for large state spaces or when probabilities are close to zero. 3. The approach assumes a clear hierarchical structure among concepts, but real-world concepts often have complex, overlapping relationships that may not fit neatly into a DAG structure. 4. The theory doesn't address how to handle noise or uncertainty in the observed data, which could significantly impact the identification of discrete states and the overall graph structure. 5. While the connection to latent diffusion models is interesting, the paper doesn't provide a concrete mechanism to leverage the theoretical insights for improving diffusion model architectures or training procedures.

Questions

1. How robust is the identification process to small violations of the invertibility condition (Condition 3.3-ii)? Are there relaxations of this condition that could make the method more applicable to real-world data while maintaining identifiability? 2. Your interpretation of latent diffusion models suggests a correspondence between diffusion steps and concept hierarchy levels. How might this insight be used to design a diffusion process that explicitly learns and respects a given hierarchical concept structure? 3. The theory assumes discrete latent variables, but the latent space in diffusion models is continuous. How do you reconcile this discrepancy, and could your framework be extended to handle continuous latent variables with discrete-like behavior?

Rating

6

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors adequately discuss the limitations of their work.

Reviewer ZL2x7/10 · confidence 1/52024-07-13

Summary

This paper presents a theoretical framework for learning discrete concepts from high-dimensional data using latent hierarchical models. The authors propose formalizing concepts as discrete latent causal variables within a hierarchical causal model, and discuss under which condition the identification of these concepts and their relationships is possible. The theoretical contributions identifying those conditions and providing both theoretical insights and empirical validation using synthetic data, along with an interpretation of latent diffusion models through the proposed framework.

Strengths

First, I apologize to the authors as I am not at all an expert in causal inference, and highly unsure about my remark (whether there be positive or negative), I also would like to mention to the author that I ensure myself the AC is aware of this. Nevertheless, concerning the strength that I identified: 1. Up to my knowledge, the paper introduces an interesting formalization of what diffusion model are doing: learning concepts over hierarchical models, which is an interesting viewpoint and could help us better understand those models. 2. The identification conditions and theorems seems well-formulated (again, not an expert). 3. The authors try to validate their theoretical claims with real data experiments, I especially like the 6.1 which seems to partially verify their claim

Weaknesses

Nevertheless, according to me (again, not an expert) this paper has problems, some more important than others. So I will separate them into major problems (**M**) and minor problems (_m_). I want to make it clear that for me, all these problems are solvable and do not detract from the quality of the paper. Let's start with what I think are the Major problems (**M**): **M1**. Practicality of Recovering Hierarchical Graphs: - I am left wanting more; the theoretical framework seems solid, but I would like to see concrete results. For example, can you recover the concept tree for the dog class in Stable Diffusion (or smaller diffusion model) ? Demonstrating this would undeniably highlight the paper's value and lead to a clear acceptance from my side. As it stands, I wonder if this framework could eventually teach us anything about diffusion models. **M2**. Empirical Evidence for Real-world Data: - The paper partially validates its claims using synthetic data. Even Section 6.1 is great, but clearly not enough to validate your claim. I'd say Figure A.4 in the appendix is another proof, and I expected maybe a comparison of this sparsity level with a real hierarchical model. Could we recover any information from the tree using this sparsity curve? **M3**. Realism of Condition 3.3: - There are doubts about the practicality and realism of Condition 3.3, especially 3.3-ii. It would be beneficial to discuss how realistic these conditions are in practical scenarios and provide more context or examples to substantiate them. Now for the minor problems: _m1_. Partial Literature Review: - The related work section doesn't discuss concept extraction, which is a significant field with many papers every year at this conference. To me, you should at least cite or mention this literature. _m2_. Clarity of Theoretical Explanations: - Some parts of the theoretical explanations, particularly in Sections 3.2 and 3.3, are dense and may be difficult for readers to follow. Additional clarifications and examples would improve the accessibility of these sections.

Questions

- Given my limited expertise, I am curious about the applicability of this framework in a supervised setting, particularly in relation to classification tasks. Could this framework be adapted, to show for example that supervised model learn only a part of the hierarchical graph relevant to classification problems?

Rating

7

Confidence

1

Soundness

3

Presentation

3

Contribution

2

Limitations

Yes, the limitations identified by the authors are accurate and well-documented. Regarding the weakness I mentioned, I reserve the right to increase the score if the authors adequately address my major concerns.

Authorsrebuttal2024-08-12

Dear Reviewer ZL2x, As the rebuttal deadline approaches, we are wondering whether our responses have properly addressed your concerns? Your feedback would be extremely helpful to us. If you have further comments or questions, we hope for the opportunity to respond to them. Many thanks, 7636 Authors

Authorsrebuttal2024-08-12

Dear Reviewer 7auw, As the rebuttal deadline approaches, we are wondering whether our responses have properly addressed your concerns? Your feedback would be extremely helpful to us. If you have further comments or questions, we hope for the opportunity to respond to them. Many thanks, 7636 Authors

Authorsrebuttal2024-08-12

Dear Reviewer ae89, As the rebuttal deadline approaches, we are wondering whether our responses have properly addressed your concerns? Your feedback would be extremely helpful to us. If you have further comments or questions, we hope for the opportunity to respond to them. Many thanks, 7636 Authors

Reviewer ae892024-08-13

I thank the authors for their diligent answer. I'm sure the changes mentioned in the rebuttal are going to improve the manuscript quite a bit. I'm going to keep my score as is.

Authorsrebuttal2024-08-12

Dear Reviewer BR1s, As the rebuttal deadline approaches, we are wondering whether our responses have properly addressed your concerns? Your feedback would be extremely helpful to us. If you have further comments or questions, we hope for the opportunity to respond to them. Many thanks, 7636 Authors

Reviewer BR1s2024-08-12

Thank the author for the response, and I am sorry for the late reply. I read through the comments, and most of my concerns are addressed, so I tend to keep my accept score.

Reviewer ZL2x2024-08-12

Thank you for the detailed and thoughtful responses to my comments. I appreciate the additional experiments and explanations you've provided, especially the practical applications (M1, M2) of your hierarchical model interpretation and the discussion of Condition 3.3. For m1, great, it's excellent, I was also thinking about concept extraction (ACE, CRAFT, ICE) which I think share some motivation with your work. Overall, I want to mention that your work has given me a lot to think about lately, and I find it very intriguing. Thank you. Given this, I’m raising my score to 7. I am still not an expert in causal inference, but I believe this work is interesting, especially for the XAI community. Good luck with the acceptance, and thank you again for this work!

Authorsrebuttal2024-08-12

Thank you so much for your kind and encouraging words – we are truly grateful for your positive and constructive feedback! It is rewarding to know that our research has provided you with new ideas to consider and we do hope our work will contribute meaningfully to the XAI community. Thank you once again for your thoughtful review and best wishes, and we will continue to explore this exciting direction!

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC