Summary
The authors propose a method for the posterior inference of DAG structure *and* function parameters with potential applicability to arbitrary functional relations between nodes. The authors modify a novel characterization of DAGs, and interpret this characterization in terms of a sorting operation which can be relaxed to allow differentiability. The authors define priors on DAGs (in the alternative space) and function parameters and based on a specific model choice characterize likelihood. They use the resulting joint distribution to iteratively sample some parameters and conduct variational inference re. others. The authors examine the performance of their proposed methodology on various synthetic and real datasets.
Strengths
- The paper is very well written. It presents the previous work, motivation for current research, and reasoning behind methodological choices very clearly.
- The paper utilizes recent, previous research intelligently and presents concrete innovations to solve well-defined problems.
- Posterior inference in the DAG structure and parameter space without some of the limitations of previous work is valuable and is likely to inspire future work.
Weaknesses
- DAG model selection results have causal implications given specific model assumptions regarding generative model of the data. ANM is such a model assumption. However, it is unclear whether the identifiability results still apply in this case, given the priors defined on DAG structure and function parameters. I think the authors' work still would be valuable as only a DAG inference method; however, since the authors present their proposal as a causal discovery + inference method, this point needs further discussion.
- I think the authors' presentation should be modified to make sure their inference method is more clearly understood. Given their initial presentation, including "posterior sampling" in the title, and frequent reference to Gibbs sampling throughout the text, leads the reader think that the authors will present results with a correct MCMC algorithm and produce a full posterior distribution. However, most promising results presented by authors include their iterative algorithm that samples from the posterior of some parameters and uses variational inference for others. This is fine as a methodological choice, but their presentation leads the reader to have higher expectations, which can become crucial depending on the use case of the reader.
- Causal sufficiency assumption prevents using the current method in problems where unobserved confounding is likely. In my opinion this is acceptable given the difficulty of the problem.
Questions
- Are there any potential difficulties with using the authors' method with other SCM model assumptions / likelihoods?
- What are grounds for baseline selection in experiments? I think this would be an important addition to the final text.
- In 6.4, were there model misspecification in other methods as well?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
I think the authors adequately address the limitations of their work overall, however see Weaknesses section above for some important caveats.