A Cross-Moment Approach for Causal Effect Estimation

We consider the problem of estimating the causal effect of a treatment on an outcome in linear structural causal models (SCM) with latent confounders when we have access to a single proxy variable. Several methods (such as difference-in-difference (DiD) estimator or negative outcome control) have been proposed in this setting in the literature. However, these approaches require either restrictive assumptions on the data generating model or having access to at least two proxy variables. We propose a method to estimate the causal effect using cross moments between the treatment, the outcome, and the proxy variable. In particular, we show that the causal effect can be identified with simple arithmetic operations on the cross moments if the latent confounder in linear SCM is non-Gaussian. In this setting, DiD estimator provides an unbiased estimate only in the special case where the latent confounder has exactly the same direct causal effects on the outcomes in the pre-treatment and post-treatment phases. This translates to the common trend assumption in DiD, which we effectively relax. Additionally, we provide an impossibility result that shows the causal effect cannot be identified if the observational distribution over the treatment, the outcome, and the proxy is jointly Gaussian. Our experiments on both synthetic and real-world datasets showcase the effectiveness of the proposed approach in estimating the causal effect.

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

Peer review

Reviewer XfVt6/10 · confidence 3/52023-06-26

Summary

This work proposes a cross-moment approach to estimating the average causal effect with latent confounders in linear SCM. One proxy variable of the latent confounder can be observed. In contrast to prior research (e.g., difference-in-difference) that requires stringent assumptions, this work shows that the causal effect can be identified and estimated using cross moments between the treatment, the outcome, and the proxy variable. It also discusses when the effect with latent confounder cannot be identified. Experiments on both synthetic and real-world data show its effectiveness.

Strengths

1. This work introduces simple arithmetic operations on the cross moments to estimate causal effects with latent confounders in linear SCM. It addresses a conventional challenge in the field of solving an OICA problem and biased estimation of DiD, which may result in bad local optima. 2. It is technically sound and the idea is clearly and concisely described. 3. It can have significance in the field and practical importance given the prevalence of the studied problem.

Weaknesses

1. The major concern is the evaluation. The baselines included are quite weak and old, e.g., KP14 published in 2014. Why the OICA method [SGKZ20] has very poor performance is not clear. Other baselines may include proximal causal inference [1], for example. Also, for real-world data experiments, other baselines are not included. And the differences between the two methods when x is not included are not explained. 2. The limitations of the proposed approach are not discussed. 3. How applicable this method is is not clear. [1] Mastouri, A., Zhu, Y., Gultchin, L., Korba, A., Silva, R., Kusner, M., ... & Muandet, K. (2021, July). Proximal causal learning with kernels: Two-stage estimation and moment restriction. In International Conference on Machine Learning (pp. 7512-7523). PMLR.

Questions

1. Why SGKZ20 is very poor and what makes KP14 and the proposed approach much better than it? 2. What are the limitations of the proposed approach? 3. What are the potential applications of the proposed method? I acknowledge I read the authors's response and I keep my positive score.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

No limitations are discussed. Consider shortening Related Work and add limitations in the last section.

Reviewer 5Myy6/10 · confidence 4/52023-07-05

Summary

This work focuses on estimating the causal effect in the presence of an unmeasured confounder within a linear causal model. The authors demonstrate that the desired causal effect can be identified by employing a single proxy variable, leveraging its non-Gaussian characteristics. Additionally, they propose a Cross-Moment algorithm for estimating the model. Furthermore, they demonstrate the effectiveness of the proposed method through experiments on both synthetic and real-world datasets.

Strengths

Estimating the causal effect becomes challenging in presence of unmeasured confounders. The proposed sufficient identification condition of this paper is novel. The analysis in this paper is presented in a logical manner. This paper is clearly written.

Weaknesses

The model is restricted as a linear causal model. The sufficient identification condition is applicable only to a single unmeasured confounder, and the proposed method may not be suitable for cases involving multiple latent variables.

Questions

Identification: 1. Is the causal effect \beta identifiable under the Non-Gaussianity assumption? This is not clear to me. In my opinion, the reason why Assumptions 2 and 3 are introduced instead of non-gaussian assumption is because one need to use the cross-moment. Am I correct? 2. If Z directly affects D, the identification of the causal effect of D on Y may be affected. It would be helpful to investigate and discuss the potential implications of this scenario in the paper. 3. In the current setting, the paper assumes the existence of a single unmeasured confounder U. However, it is worth exploring and addressing the situation where there are multiple unmeasured confounders U. This could enhance the comprehensiveness and applicability of the proposed method. Related work: The following paper may be related to this work and is deserved to discuss. Shuai, K., Luo, S., Zhang, Y., Xie, F., & He, Y. "Identification and Estimation of Causal Effects Using non-Gaussianity and Auxiliary Covariates." arXiv preprint arXiv:2304.14895 (2023).

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Estimating the causal effect using the proposed method requires the availability of one proxy variable that is associated with the unmeasured confounder but does not directly affect the treatment variable. However, in practical scenarios, it can be challenging to identify or obtain such a proxy variable.

Reviewer p1qb5/10 · confidence 3/52023-07-07

Summary

This paper introduces an innovative technique for estimating the causal effect of a treatment on an outcome within linear structural causal models. This method utilizes cross moments, which are statistical moments derived from the joint distribution of the treatment and outcome variables, to quantify the causal effect. The authors demonstrate that this approach can relax the conventional assumption of a common trend in the difference-in-difference estimator, allowing for causal effect estimation in scenarios where traditional methods may fall short. To validate the effectiveness of the proposed method, the authors provide both simulation studies and a real-world application. These empirical analyses showcase the promising potential of this novel approach for estimating causal effects in linear structural causal models.

Strengths

* Novelty: The paper introduces an innovative approach to estimating causal effects in linear structural causal models with latent confounders by leveraging cross moments. This method deviates from conventional approaches and exhibits the potential to yield more precise estimates within specific contexts. * Rigor: The authors establish a rigorous theoretical framework for their proposed method, delineating the conditions that allow for the identification of the causal effect and the applicability of the method. Furthermore, they substantiate their claims through comprehensive simulation studies and a real-world application, bolstering the robustness and effectiveness of the approach. * Significance: Estimating causal effects is a crucial task across various domains, and the proposed method holds substantial importance as it can potentially deliver more accurate estimates in specific scenarios. By addressing the limitations of traditional approaches, this method offers a valuable contribution to the field of causal effect estimation.

Weaknesses

The authors should compare their proposed method to more existing methods for estimating causal effects, such as negative outcome control. This will help readers understand how the proposed method compares to existing methods in terms of accuracy and efficiency.

Questions

See weakness

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Discussed.

Reviewer dXHi5/10 · confidence 4/52023-07-18

Summary

The authors consider the estimation of causal effect in linear SCM with independent errors when there is a latent confounder U and one proxy variable of U (negative control outcome). They generalize the DiD literature by relaxing the assumption of common trends and propose a general identification formula under the stated structural assumptions. The authors show that under some restrictions on the latent confounder moments and on the U-Z, U-D relations, causal effects are uniquely identified nonparametrically using the identification formula. Furthermore, they propose a general estimation algorithm based on the cross-moments of Z and D. In addition, the authors provide an ``impossibility result" which shows that under fully gaussian linear SCM, causal effects can not be uniquely identified. They illustrate their proposed method in simulation and a data example.

Strengths

Correcting the bias due to residual confounding using proxy variables (negative controls) is an emerging topic in causal inference. The authors consider the more difficult task that seeks to correct the bias with only one proxy variable. Under the assumed linear structural model with independent errors, the authors utilized a well-known identity that can be thus used as an identification formula. Theorem 1, which provides the uniqueness guarantees, is novel and motivates a computationally easy estimation algorithm, which is an improvement in comparison to other proximal learning methods. The formal results are rigorous and nontrivial. The theoretical guarantees provide a meaningful illustration of the limitations of the proposed identification formula. The paper is well-written and easy to follow.

Weaknesses

The authors provide results only for linear SCM with independent errors. Both linearity and exogenous errors are fairly strong assumptions that are not likely to hold in practice. Moreover, the theoretical results heavily depend on both assumptions and are not likely to extend to other SCMs. On line 117, D is assumed to be a binary treatment. On line 119, the authors explicitly define the causal estimand of interest as the average causal effects on the treated. However, the SCM (line 133) states that $D = \alpha_dU +\varepsilon_d$, which, coupled with the assumption of independent zero mean errors (line 134), yields that $$\Pr(D=1)=E[D]=\alpha_dE[U] + E[\varepsilon_d]=0$$ That is, if D is binary, the SCM implies that it is a deterministic random variable that equals 0 with probability one. In addition, in Algorithm 1, $num$ is identically the same for all $n$ whenever $D$ is binary. The authors are most likely well aware of this inconsistency since in the simulation study $D$ is not taken to be a binary variable. Their proposed method works well for non-binary D, but causal estimands should be adjusted accordingly. Section 3.2 (lines 209-228) is well known in the literature (see for example the recent review by Roth et al. 2023 ``What’s trending in difference-in-differences? A synthesis of the recent econometrics literature"). Experiments under misspecification (linearity, additional latent variables, etc.) are not presented. The robustness of the proposed methods is an open question.

Questions

In the data example, estimation using the cross-moments algorithm is performed on the residuals of the outcom~covariates regression. Are there any theoretical guarantees (e.g., similar to Theorem 1) when covariates are included? are the covariates also need to have a linear relation to the treatment/outcome/negative control for the uniqueness of $\beta$? Proximal learning (Tchegen Tchegen et al, cited by the authors) provides a flexible approach for estimating causal effects with latent confounders when there are at least two proxy variables. In practice, many studies do have more than one possible proxy variable. Do you think it is possible to extend the cross-moments algorithm for scenarios with more than one negative control (e.g., under linear SCM)?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

As already stated, the theoretical results strongly rely on the linearity and independent error assumptions. The authors did not adequately address those limitations.

Reviewer XfVt2023-08-11

I thank the authors for answering my questions and doubts. The authors have especially agreed to include the explanation for the bad performance of the method in [SGKZ20] in the next version and also presented addtional results to show their algorithm can outperform more recent models. I believe by including both will improve the paper, and hence I change my score to "Weak Accept". However, it is still not clear to me why the compared methods are different for synthetic data and real-world data.

Authorsrebuttal2023-08-11

Thanks for reading the responses. We will add what was suggested by the reviewer in the revised paper. Regarding your question, please note that in the real dataset, we only have access to one proxy variable (employment level before the rise in the minimum wage). The methods in [KP14] or [TYC+20] require at least two proxy variables. Moreover, it is unclear which one of the covariates can be served as proxy variable W in these works.

Reviewer 5Myy2023-08-12

Regarding question 2

Thanks for your response! Regarding your example in Q2, what if we assume the faithfulness assumption, can we identify those two models from observed variables? In my view, the reason why the causal effect of $D$ on $Y$ can be uniquely identified is that the observed descendants of $U$ are not the same as the descendants of $D$. Please correct me if I'm wrong.

Authorsrebuttal2023-08-13

Thanks for reading the responses. Under the faithfulness assumption, if the observational distribution is generated based on the original causal graph (with just an edge from $Z$ to $D$), the second model that we proposed in the response violates faithfulness assumption as $Z$ and $Y$ should be d-separated given $D$ and $U$ which is not the case in the causal graph of this model. Thus, under the faithfulness assumption, the original model is uniquely identifiable. Regarding uniquely recovering $\beta$, as mentioned by the reviewer, the latent confounder $U$ does not have the same observed descendants as $D$. Otherwise, we can swap their corresponding exogenous noises and get a new model with a different causal effect similar to the example in Section 4.1 of [SGKZ20].

Reviewer 5Myy2023-08-16

Thanks for your response

Thanks for your response. The authors addressed my concern, so I will keep my score.

Reviewer dXHi2023-08-16

Thank you for the response. The authors addressed my questions and concerns.

Reviewer p1qb2023-08-18

Thanks for the reply

The author successfully addressed my questions in their rebuttal stage, and I would like to keep my score to vote for acceptant.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC