Re: Official comment
Thank you very much for the reply. We hope the following addresses the remaining follow-up questions.
**Eq. 1: So are you saying that the RHS of Eq. 1 is an estimator for the LHS? ...**
Apologies, but we are genuinely confused where this conclusion is coming from. We don't see anything in our rebuttal that suggests that, and it would be helpful to have the quote of the passage leading to it. Eq. 1 is indeed a "a statement about the true data generating process".
**Fig. 1: ...in your rebuttal you say that confounding between action variables is allowed, but not preceding action variables causing future actions...**
Again, we don't see anywhere in our rebuttal a statement that $D$ variables cannot be causing others. In fact, we explicitly say the opposite ("*...nothing in our results change at all if ... past $D$ variables pointing to all future $D$ variables.*"). Perhaps it's useful to recall the difference between interventions and random variables within our context.
In a fully connected DAG according to ordering $(D_1, X_1, D_2, X_2)$ where $D_1$ and $D_2$ are random instead of controlled, we have the edges
$D_1 \rightarrow X_1, D_1 \rightarrow X_2, D_1 \rightarrow D_2$, $X_1 \rightarrow D_2$, $X_1 \rightarrow X_2$, $D_2 \rightarrow X_2$.
When $D_1$ is controlled to $d_1$ and $D_2$ to $d_2$, one graphical characterization is
$d_1 \rightarrow X_1, d_1 \rightarrow X_2$, $X_1 \rightarrow X_2$, $d_2 \rightarrow X_2$,
where lower case here indicates that we are talking about exogenous fixed indices (squares in Fig. 1).
There isn't anything else to be said if $d_2$ is functionally independent of the past, but even then this is not at all an issue. This is because whether $d_2$ is a function or not of $(d_1, X_1)$, it won't affect identification strategies such as the sequential back-door/g-formula see e.g. chapter 4 of [32]. We can still carry $d_2$ symbolically into any averaging over the past, even if $d_2$ is functionally related to $(D_1, X_1)$ and we treat $D_1$ as uncontrolled and average over it (since we don't average over the past, this point is redundant anyway). Other representations such as SWIGs suggest adding edges between fixed indices with functional dependencies see e.g. Chapter 19 of [22], but if we were to adopt SWIGs in Fig. 1 the diagram would become an incomprehensible mess. Technically speaking, even edges like $d_1 \rightarrow X_1$ are "unnecessary", as the graphical model is meant to represent the independence structure of a distribution over random variables, and $d_1$ isn't one (preserving "$d_1 \rightarrow \dots$") saves us from having to label the nodes as $X_1(d_1)$ etc.).
To summarize, our Fig. 1 is a cosmetic choice among other plausible choices and exemplifies already the case of non-dynamic regimes with no ambiguity. There wasn't really much of a deep point we were trying to make (other than not requiring any Markovian assumptions connecting past and future), and we are earnestly surprised that this is raising a discussion...
**Why did you make the assumption of randomization in the first place?**
This was just a way of saying that we assume to have access to the distribution of (single-world) potential outcomes, where our predictions lie. Whether we obtain it by (say) controlled experiments, sequential ignorability, proxies, instrumental variables etc. is orthogonal to our main results. We thought readers would appreciate if we focused on the main novel aspects of our contribution. We are happy to make this point more explicitly in the introduction.
We really appreciate your engagement, and we hope the above has been helpful.