Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data

Discovering causal relationship using multivariate functional data has received a significant amount of attention very recently. In this article, we introduce a functional linear structural equation model for causal structure learning when the underlying graph involving the multivariate functions may have cycles. To enhance interpretability, our model involves a low-dimensional causal embedded space such that all the relevant causal information in the multivariate functional data is preserved in this lower-dimensional subspace. We prove that the proposed model is causally identifiable under standard assumptions that are often made in the causal discovery literature. To carry out inference of our model, we develop a fully Bayesian framework with suitable prior specifications and uncertainty quantification through posterior summaries. We illustrate the superior performance of our method over existing methods in terms of causal graph estimation through extensive simulation studies. We also demonstrate the proposed method using a brain EEG dataset.

Paper

Similar papers

Peer review

Reviewer Y9Hq5/10 · confidence 3/52023-06-19

Summary

The paper develops a causal discovery method for directed cyclic graph. The proposed method is a two-step approach which utilizes a Bayesian approach to reduce the dimension of functional data and performs the causal structure learning on the learned embeddings. Experiments on the synthetic and real-world datasets well demonstrate the effectiveness of the proposed method.

Strengths

1. There are seldom paper studying the causal discovery method on the directed cyclic graph. However, in some particular applications, there may exist cycles in the causal graphs. 2. The proposed method outperforms baselines in various settings. 3. The results of the brain EEG data provide some insights on the organization of brain regions.

Weaknesses

1. It lacks some classical and SOTA baselines, e.g., NOTEARS[1], DAG-GNN[2], CSIvA[3] [1] Zheng X, Aragam B, Ravikumar P K, et al. Dags with no tears: Continuous optimization for structure learning[J]. Advances in neural information processing systems, 2018, 31. [2] Yu Y, Chen J, Gao T, et al. DAG-GNN: DAG structure learning with graph neural networks[C]//International Conference on Machine Learning. PMLR, 2019: 7154-7163. [3] Learning to induce causal structure. ICLR 2023. 2. It is hard to set so many priors for the proposed model in practice. 3. The presentation of the manuscript could be improved. Fig.1 is hard to understand. 4. For the results of real data, it is better to also show the results of baselines.

Questions

Please refer to weaknesses.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

1 poor

Contribution

2 fair

Limitations

No I have not found the limitations and broader societal impacts of this manuscript.

Reviewer 3JSM6/10 · confidence 3/52023-06-28

Summary

This paper discusses the inference of causal relationships between multivariate functions through the proposal of a causal discovery model. The model involves embedding the functional nodes into a lower-dimensional space and separating the structural equation model into two components: the projection onto the space and its orthogonal complement in the Hilbert space defined on the domain. The identifiability of the proposed model is proven based on several assumptions, including disjoint cycles, as described in the Supplementary Materials. The paper presents a model inference method utilizing a fully Bayesian approach. Experimental results using simulated data demonstrate that the proposed method outperforms both conventional causal discovery methods for multivariable functional data and other conventional causal discovery methods. Additionally, insightful observations are made from the application of the proposed method to brain EEG data.

Strengths

The strongest aspect of this paper is its introduction of a causal discovery model for multivariate functional data, which accommodates the presence of cycles. This is a significant contribution since many multivariate functional datasets, such as EEG data, inherently involve cycles, including feedback loops. Additionally, the paper proves the identifiability of the model and presents a fully Bayesian approach to model inference. The effectiveness of the proposed method is demonstrated through the utilization of both simulated data and real-world EEG data.

Weaknesses

The discussion of the results of the brain EEG data is inadequate. While there are differences in connectivity between the alcoholic and control groups, readers are unable to determine the significance of these findings because most readers may not possess a background in brain science. There may be a minor mistake: from my understanding, the subscript $j=1$ below the superscript $m_j$ in line 116 might actually be $u=1$.

Questions

Is it possible to consider providing additional discussion on the results of the brain EEG data, in order to help readers ascertain the validity of the findings and comprehend their significance?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The limitation of the proposed method lies in the assumptions required to establish the identifiability of the model.

Reviewer nZkc7/10 · confidence 3/52023-07-05

Summary

The authors propose a causal model, i.e. a linear structural equation model for multivariate functional data. The authors' proposed model does not require the dependency structure to conform to a DAG, allowing for modeling cyclic cause-effect relationships. Given certain assumptions on exogenous variables and cycle structure, the authors show the identifiability of their model; a central assumption of which is that the causal interrelations between the variables pertain only to a subset of the function space which the observations inhabit. The authors propose an MCMC sampling procedure to sample from the posterior of the parameters given prior hyperparameters. The authors test their model under various conditions in synthetic data, in addition to a real world application.

Strengths

- The authors' research is soundly motivated, the central observation regarding the importance of being able to handle cyclic relationships in functional data is well-justified, especially given the application areas they consider. - The authors' central modeling assumption, that the causal interrelations pertain to a subset of the function space is interesting and has the potential to be utilized and expanded upon future research. - Minus some limiting assumptions, the authors' identifiability results and inference procedure is valuable. - The authors provide a rigorous exploration of various aspects of their model in comparison with baselines in synthetic data.

Weaknesses

- The authors make a series of assumptions (1-6) in order to achieve identifiability, with the first one being causal sufficiency. Both the assumption of causal sufficiency and making simplifying assumptions in general is a widely used practice especially when expanding causal analysis to previously unexplored territory, and is usually acceptable to expedite the initial analysis of a new idea. Although it is still understandable that the authors make this assumption, for the kinds of data they are hoping to conduct causal analysis on this assumption unfortunately seems to be especially likely to be violated. For example, it sounds very improbable for measurements made in specific parts of the brain to not have latent confounding. A similar case against disjoint cycles can also be made, while the rest of the assumptions seem less offensive. That being said, I think the authors' contributions justify these limitations. - Given the presence of these limiting assumptions, some external validation through real data experiments is called for. The authors' experiment with real data present no external validation (e.g. discovering some known causal relationships among various brain regions).

Questions

- In addition to the aforementioned assumptions, using a linear model is another disputable choice, as the authors note. Did the authors experiment with nonlinear generated data to examine the effects of this potential source of model misspecification? - Brain imaging is an obvious (and worthy) example for the utilization of such a modeling approach. Are there any other impactful application fields that can benefit from such work? - In L99, is it supposed to be $\mathcal{H}_\ell$ to $\mathcal{H}_j$? - In L116, is it supposed to be $\\{(t_{ju}, X_{ju})\\}^{m_j}_{u=1}$?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors generally transparent about their modeling assumptions and potential effects thereof. However some of the points I mentioned above would benefit from more elaboration in the final paper.

Reviewer 236D5/10 · confidence 4/52023-07-06

Summary

In this paper, the author proposes an operator-based non-recursive linear structural equation based novel causal discovery framework that identifies causal relationships among functional objects in the presence of cycles and additional sampling noises. Furthermore, author demonstrates the effectiveness of proposed method from experiment and theory.

Strengths

The motivation is clear and the writing is well. The experiment results show that the proposed method significantly outperforms the baselines. The causal identifiability proof is impressive.

Weaknesses

Please refer to the Questions.

Questions

1. The proposed year of your baselines fLiNG, LiNGAM, PC, CCD are 2022, 2006, 1991, 1996 respectively, more recent baselines should be considered. According to line 282 and 283, the codes of Lee (2022) and Yang (2022) are not available. However, fLiNG’s code is also not public (I tried my best to find it but failed). Hence, how do you obtain the results reported in the paper ? 2. In line 18, you said casual discovery is popular in machine learning. But both in introduction and related work, I do not find relative description to support this standpoint. 3. How to obtain equation (3) ? It appears to be left-multiplying Equation (1) by and respectively. If so, the term in first row should be . 4. In line 116, Why does j vary from 1 to ? The range of j should be from 1 to p. 5. Many symbols are used without definition, like Beta in line 241, Dir and IG in line 254. Some symbols are easy to search for their meanings, while others are not as straightforward. 6. The equation (7) are not mentioned in Section 3, is it really necessary ? 7. The name of Section 3 is Model inference. However, nothing but the prior distribution of parameters are introduced, this is confusing. 8. Reference [Fangting Zhou, Kejun He, and Yang Ni. Causal discovery with heterogeneous observational data] mentions that they do not restrict their model to be acyclic, which is in conflict with your statement in line 56 (Related work section).

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

No apparent negative societal impacts.

Reviewer ucHi7/10 · confidence 4/52023-07-12

Summary

The paper considers the problem of learning causal structure from multivariate functional data that may involve cyclic interactions. It employs an adaptive mapping to a lower dimensional space that retains the relevant causal information, in combination with a Bayesian framework that obtains posterior estimates through MCMC sampling. Effectiveness of the model is compared against several alternatives on synthetic and real-world data.

Strengths

The paper is exceptionally well written and contains a comprehensive and powerful approach to cyclic causal discovery. Many of the steps involved are not new by themselves, but brought together effectively to produce an elegant and novel method that shows promising performance. The descriptions and theoretical derivations are concise, clear and consistent, and each of the steps involved follows logically to reduce a challenging problem to an effective, fully Bayesian solution. I particularly liked the way the adaptive basis expansion in 3.2 was incorporated into the overall method. Assumptions are strong, but clearly and carefully spelled out, leading to the desired identifiability conclusion in Thm2.1. The synthetic data experimental evaluation is convincing (although due to the required assumptions inevitably biased in favour of the proposed approach), and the alcoholic EEG application is interesting (though hard to evaluate for non-experts).

Weaknesses

- although I really like the paper, I think the main contribution lies mostly in the way the various steps are brought together in a clear and coherent way than in fundamentally new insights or approaches. (Nevertheless, I do consider the end result a valuable and novel contribution.) - assumptions are quite strong: to (still) see causal sufficiency feature so prominently is disappointing, the ‘disjoint cycles’ should be unnecessary (as admitted by the authors), and ‘non-Gaussianity’ may seem generic but implies identifiability may rely on weak distributional signals (data hungry) and it excludes the challenging complications encountered in the linear Gaussian case that e.g. CCD was specifically designed to handle. - one confusing aspect (at first) was that in the beginning of the paper the introduction of the time measurements seemed to suggest we are doing time-series analysis, but on closer inspection it is actually closer to standard observational causal inference (with cycles), right? - some questionable claims, e.g. l.245 ‘strongly prevents false discoveries’ could equally be stated as ‘heavily biased towards sparsity’, and l.307 ‘strong evidence of the effectiveness of FENCE compared to existing methods’ should contain the caveat ‘under the stated assumptions 1-6’. In particular, starting from PCA for LINGAM / CCD seems questionable, where LiNGAM is designed for DAGs, and independence based methods like CCD do not (need to) rely on functional assumptions that are essential for FENCE.

Questions

- what is the relation/difference between your model assumptions and e.g. the simple SCMs in Bongers et al. [‘Foundations of structural causal models with cycles and latent variables.’,2018]? - how do you determine Kj for the lower dimensional space embedding? - is it true that you essentially treat multiple observations at different time points as independent observations on different system instances in equilibrium? - in. Table 1, how do you compare non-invariant edges in the equivalence class of e.g. PC/CCD vs. the ground truth? (as ‘half wrong’?) - given the causal sufficiency assumption, how come you find bidirected edges in Fig.2? Other remarks (for lack of a better place to put them): - title seems grammatically a bit weird - 5 ‘enhance interpretability’ => I understand the practical effectiveness in going via a lower dimensional internal representation, but the output does not reflect this, right? So how does it affect ‘interpretability’? - 88 ‘cyclic component’ => this definition seems a bit off as it allows for vertices in a subgraph that are not part of a cycle (so only the full graph G can be a ‘maximal’ cyclic component). I would suggest using the standard term ‘strongly connected component’. - 258 ‘simulate posterior samples through MCMC’ => actually I would not mind seeing a bit more on this step in the main paper

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

NA

Reviewer 236D2023-08-11

Comments after Rebuttal

Thanks for addressing my issues, Although there are still certain defect to be improved in this paper, the motivation is very interesting, and the proposed behaviour is technical solid. Concretely, I prefer to change my rating.

Reviewer Y9Hq2023-08-14

The authors have properly solved my concerns. Hence I raise my score accordingly.

Reviewer nZkc2023-08-15

Thanks for the response

I thank the authors for their response and recommend the acceptance of the paper. I think integrating their responses regarding model misspecification, modeling assumptions, and how their results corroborate existing findings into the final version of the paper will further improve their work.

Authorsrebuttal2023-08-15

Thank you for your reply. We are glad that our response helps. We will integrate our responses regarding misspecification analysis, model assumptions and the validation of our real data analysis results in the revised version of the paper.

Reviewer ucHi2023-08-18

The other reviews and author rebuttals have strengthened my original impression that this is an interesting and worthwhile contribution that deserves to be in the conference. Hence I will stay with 'Accept'.

Reviewer 3JSM2023-08-21

Thank you for your rebuttal!

Thank you very much for your answer to my question! I understood the validity of the experimental results on the brain EEG data. I re-think my opinion by considering the reviews of other reviewers, I decide to change my rating. I sincerely apologize for my late response.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC