Summary
The authors tackle the problem of decomposing the source of variation in causal analysis. In many real-world problems, there are many factors that can introduce spurious correlation between a treatment and outcome, and breaking down the contribution of each variable to that spurious correlation can be useful in many fields where an explanation is important. The authors propose a methodology for both Markovian and semi-Markovian models, based on the 'abduction-action-prediction' method. It allows for the application of evidence to only a subset of exogenous variables, which then allows us to separate the variation from each exogenous variable.
Strengths
This paper is well-written and well-motivated. The problem it seeks to address is important, and the introduction does a good job at setting the theoretical foundation and describing the problem. For the most part, the terminology is introduced at a good pace, making the overall narrative easier to follow. I appreciate the authors interspersing examples with figures, since without them, the paper would likely bog down with all the math. Overall, I found this a compelling approach to an important problem.
Weaknesses
I think, for such an equation-heavy paper, the notation could be made clearer/more explicit in sections. For example, Proposition 1 is the first time that we see the $e^{U_1}$ notation, which it seems like means 'this evidence is applied to everything except for $U_1$'. That's a fairly counter-intuitive notation for such a thing, since generally a subscript or superscript implies that we're using those variables in someway, rather than explicitly excluding them (I realize you need to use them to exclude them, but conceptually, the notation is non-obvious for me). If I'm not misunderstanding what this notation means, then a more explicit definition before Proposition 1 would be helpful, given how central it is to the rest of the paper.
As another notation point, Theorem 2 uses the notation $z_{-[i]}$, but until now, the $Z_{[i]}$ notation has meant, as defined in Definition 2, "all $Z_i$'s, from 1 to i", and it's not obvious how this translates to a negative.
The COMPAS demonstration is great, but since it clearly lacks ground truth, additional synthetic experiments would be helpful to help demonstrate both the correctness of the theory and how this method can be used in practice on other types of data.
Questions
Is there a list of all of assumptions required for your method? In Section 4.1, you say that the decomposition needs to follow a topological order of the variables U_i, but are there others?
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
In terms of social impact, this method is more likely to have a positive social impact, since causal inference is often applied in domains where spurious variation is present but where the decisions made can have significant effects on people's lives (e.g., the COMPAS example)