Summary
The authors consider the estimation of causal effect in linear SCM with independent errors when there is a latent confounder U and one proxy variable of U (negative control outcome).
They generalize the DiD literature by relaxing the assumption of common trends and propose a general identification formula under the stated structural assumptions.
The authors show that under some restrictions on the latent confounder moments and on the U-Z, U-D relations, causal effects are uniquely identified nonparametrically using the identification formula. Furthermore, they propose a general estimation algorithm based on the cross-moments of Z and D.
In addition, the authors provide an ``impossibility result" which shows that under fully gaussian linear SCM, causal effects can not be uniquely identified.
They illustrate their proposed method in simulation and a data example.
Strengths
Correcting the bias due to residual confounding using proxy variables (negative controls) is an emerging topic in causal inference. The authors consider the more difficult task that seeks to correct the bias with only one proxy variable.
Under the assumed linear structural model with independent errors, the authors utilized a well-known identity that can be thus used as an identification formula. Theorem 1, which provides the uniqueness guarantees, is novel and motivates a computationally easy estimation algorithm, which is an improvement in comparison to other proximal learning methods.
The formal results are rigorous and nontrivial. The theoretical guarantees provide a meaningful illustration of the limitations of the proposed identification formula.
The paper is well-written and easy to follow.
Weaknesses
The authors provide results only for linear SCM with independent errors. Both linearity and exogenous errors are fairly strong assumptions that are not likely to hold in practice. Moreover, the theoretical results heavily depend on both assumptions and are not likely to extend to other SCMs.
On line 117, D is assumed to be a binary treatment. On line 119, the authors explicitly define the causal estimand of interest as the average causal effects on the treated. However, the SCM (line 133) states that $D = \alpha_dU +\varepsilon_d$, which, coupled with the assumption of independent zero mean errors (line 134), yields that
$$\Pr(D=1)=E[D]=\alpha_dE[U] + E[\varepsilon_d]=0$$
That is, if D is binary, the SCM implies that it is a deterministic random variable that equals 0 with probability one.
In addition, in Algorithm 1, $num$ is identically the same for all $n$ whenever $D$ is binary.
The authors are most likely well aware of this inconsistency since in the simulation study $D$ is not taken to be a binary variable.
Their proposed method works well for non-binary D, but causal estimands should be adjusted accordingly.
Section 3.2 (lines 209-228) is well known in the literature (see for example the recent review by Roth et al. 2023 ``What’s trending in difference-in-differences? A synthesis of the recent econometrics literature").
Experiments under misspecification (linearity, additional latent variables, etc.) are not presented. The robustness of the proposed methods is an open question.
Questions
In the data example, estimation using the cross-moments algorithm is performed on the residuals of the outcom~covariates regression. Are there any theoretical guarantees (e.g., similar to Theorem 1) when covariates are included? are the covariates also need to have a linear relation to the treatment/outcome/negative control for the uniqueness of $\beta$?
Proximal learning (Tchegen Tchegen et al, cited by the authors) provides a flexible approach for estimating causal effects with latent confounders when there are at least two proxy variables. In practice, many studies do have more than one possible proxy variable. Do you think it is possible to extend the cross-moments algorithm for scenarios with more than one negative control (e.g., under linear SCM)?
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
As already stated, the theoretical results strongly rely on the linearity and independent error assumptions. The authors did not adequately address those limitations.