Summary
The paper considers the problem of learning causal structure from multivariate functional data that may involve cyclic interactions.
It employs an adaptive mapping to a lower dimensional space that retains the relevant causal information, in combination with a Bayesian framework that obtains posterior estimates through MCMC sampling. Effectiveness of the model is compared against several alternatives on synthetic and real-world data.
Strengths
The paper is exceptionally well written and contains a comprehensive and powerful approach to cyclic causal discovery.
Many of the steps involved are not new by themselves, but brought together effectively to produce an elegant and novel method that shows promising performance. The descriptions and theoretical derivations are concise, clear and consistent, and each of the steps involved follows logically to reduce a challenging problem to an effective, fully Bayesian solution. I particularly liked the way the adaptive basis expansion in 3.2 was incorporated into the overall method.
Assumptions are strong, but clearly and carefully spelled out, leading to the desired identifiability conclusion in Thm2.1.
The synthetic data experimental evaluation is convincing (although due to the required assumptions inevitably biased in favour of the proposed approach), and the alcoholic EEG application is interesting (though hard to evaluate for non-experts).
Weaknesses
- although I really like the paper, I think the main contribution lies mostly in the way the various steps are brought together in a clear and coherent way than in fundamentally new insights or approaches. (Nevertheless, I do consider the end result a valuable and novel contribution.)
- assumptions are quite strong: to (still) see causal sufficiency feature so prominently is disappointing, the ‘disjoint cycles’ should be unnecessary (as admitted by the authors), and ‘non-Gaussianity’ may seem generic but implies identifiability may rely on weak distributional signals (data hungry) and it excludes the challenging complications encountered in the linear Gaussian case that e.g. CCD was specifically designed to handle.
- one confusing aspect (at first) was that in the beginning of the paper the introduction of the time measurements seemed to suggest we are doing time-series analysis, but on closer inspection it is actually closer to standard observational causal inference (with cycles), right?
- some questionable claims, e.g. l.245 ‘strongly prevents false discoveries’ could equally be stated as ‘heavily biased towards sparsity’, and l.307 ‘strong evidence of the effectiveness of FENCE compared to existing methods’ should contain the caveat ‘under the stated assumptions 1-6’. In particular, starting from PCA for LINGAM / CCD seems questionable, where LiNGAM is designed for DAGs, and independence based methods like CCD do not (need to) rely on functional assumptions that are essential for FENCE.
Questions
- what is the relation/difference between your model assumptions and e.g. the simple SCMs in Bongers et al. [‘Foundations of structural causal models with cycles and latent variables.’,2018]?
- how do you determine Kj for the lower dimensional space embedding?
- is it true that you essentially treat multiple observations at different time points as independent observations on different system instances in equilibrium?
- in. Table 1, how do you compare non-invariant edges in the equivalence class of e.g. PC/CCD vs. the ground truth? (as ‘half wrong’?)
- given the causal sufficiency assumption, how come you find bidirected edges in Fig.2?
Other remarks (for lack of a better place to put them):
- title seems grammatically a bit weird
- 5 ‘enhance interpretability’ => I understand the practical effectiveness in going via a lower dimensional internal representation, but the output does not reflect this, right? So how does it affect ‘interpretability’?
- 88 ‘cyclic component’ => this definition seems a bit off as it allows for vertices in a subgraph that are not part of a cycle (so only the full graph G can be a ‘maximal’ cyclic component). I would suggest using the standard term ‘strongly connected component’.
- 258 ‘simulate posterior samples through MCMC’ => actually I would not mind seeing a bit more on this step in the main paper
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.