A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured Confounding

In many social, behavioral, and biomedical sciences, treatment effect estimation is a crucial step in understanding the impact of an intervention, policy, or treatment. In recent years, an increasing emphasis has been placed on heterogeneity in treatment effects, leading to the development of various methods for estimating Conditional Average Treatment Effects (CATE). These approaches hinge on a crucial identifying condition of no unmeasured confounding, an assumption that is not always guaranteed in observational studies or randomized control trials with non-compliance. In this paper, we proposed a general framework for estimating CATE with a possible unmeasured confounder using Instrumental Variables. We also construct estimators that exhibit greater efficiency and robustness against various scenarios of model misspecification. The efficacy of the proposed framework is demonstrated through simulation studies and a real data example.

Paper

References (24)

Scroll for more · 12 remaining

Similar papers

Peer review

Reviewer iUwm4/10 · confidence 5/52024-07-06

Summary

In this paper, the authors proposed a general framework for estimating CATE with a possible unmeasured confounder using instrumental variables. They construct estimators that exhibit efficiency and robustness against various scenarios of model misspecification. The efficacy of the proposed framework is demonstrated through simulation studies and a real data example.

Strengths

Strengths: The robustness and efficacy of the proposed method are shown in theory and simulations.

Weaknesses

Please refer to Questions.

Questions

The authors might consider scenarios where one might be interested in heterogeneous treatment effect on $X'$, where $X'$ is a subset of $X$.

Rating

4

Confidence

5

Soundness

2

Presentation

2

Contribution

2

Limitations

Please refer to Questions.

Reviewer M1zT5/10 · confidence 2/52024-07-10

Summary

The author mainly introduces a method of Directly Learning using Instrumental Variables (IV-DL) to estimate the conditional average treatment effect (CATE) $\Delta(x)$ and optimal Individualized Treatment Regime (ITR) $\hat{d(x)}$ in the presence of unobserved confounding. They propose two efficient and robust estimators, IV-RDL1 and IV-RDL2, by residualizing the outcome. The authors conduct two simulation settings and use real-world data to demonstrate the efficiency of the approach.

Strengths

This paper primarily focuses on estimating conditional average treatment effects and determining the optimal individualized treatment regime using the Direct Learning with Instrumental Variable Approach. It presents a well-structured logical framework to discuss this concept.

Weaknesses

I think overall the authors did interesting research, but my main concerns are listed below. The variables A, Y, and Z in this paper are all binary variables. It would be beneficial to discuss how the framework of IV-DL can be extended or adapted to handle a continuous instrumental variable 𝑍. The current focus on binary variables may limit the generalizability of the findings. In Section 5, the authors do not provide a detailed introduction to work similar to IV-RDL2. A more thorough explanation of the differences and connections between this work and related research would enhance the reader's understanding of the unique contributions and context of the presented study.

Questions

1. In section 6, why is the performance of IV-RDL1 than IV-RDL2 in three metrics? 2. The authors should conduct some experiments to show whether IV-RDL1 and IV-RDL-2 are robust compared to other methods in the presence of misestimation. 3. The authors miswrite “$f(x)=\tilde{x}^{T}\mathbf{\beta}$” as “$\Delta(x)=\tilde{x}^{T}\mathbf{\beta}$” in line 152.

Rating

5

Confidence

2

Soundness

2

Presentation

3

Contribution

3

Limitations

The authors discuss some of the algorithm's shortcomings.

Reviewer qaqj7/10 · confidence 4/52024-07-12

Summary

The authors study the problem of estimating the conditional average treatment effect (CATE) under the assumption of unmeasured confounding. The authors focus on the specific scenario where some observed variable acts as instrument w.r.t. unmeasured confounder but might be confounded by some other observed confounder, so that standard IV methods may fail. They derive a method which extends Direct Learning (DL) by an additional scaling factor of the outcomes. This scaling factor is the CATE of the instrument on the treatment, which can be estimated with standard methods (e.g., DL). In a simulation study, the authors compare the proposed method to a set of baseline methods.

Strengths

The problem and method are well-presented. The resulting method is simple but elegant. It extends Direct Learning by, first estimating the CATE of the instrument on the treatment, and estimates the CATE of the treatment on the outcome through Direct Learning leveraging the result of the first step. The authors prove identifiability under a provided set of assumptions. They propose two ways to residualize the outcomes in order to reduce the variance of the estimator, and provide a set of sufficient conditions (in terms of correctly specified nuisance functions) under which the estimator yields consist CATE estimates.

Weaknesses

Assumption 2.f seems rather strong as the unobserved confounder can only additively affect the treatment. The data generating process in the experimental section violates this assumption. As the proposed method still outperforms the baselines, this may suggest that the method is less sensitive to the assumption. Though, that should be studied empirically in more detail. Compared to other recent publications focusing on estimating the (conditional) treatment effect, the assumed data generating process in the simulation seems overly simple. Other methods involving GPs, normalizing flows, and other highly non-linear models allow for high-dimensional confounders. They are typically assessed using semi-artificial data (e.g., with images as confounders and image labels as confounding mechanism). Without such hard problems and the corresponding baseline methods, it is hard to assess the overall practical value of the proposed method.

Questions

The paper is clear; no questions but some minor comment: - The term instrumental variable for Z might be a bit confusing as Z and Y are confounded; Z is an IV w.r.t. unmeasured confounder. It may help to clarify that in the very beginning. - Figure 1 might be improved in terms of order and size of the nodes; having the instrument on the right, the treatment at the top, and the outcome on the left is rather unconventional. - lines 151-164 have likely little value as this should be known to the Neurips audience

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

NeurIPS Paper Checklist is provided; no concerns.

Reviewer M5di6/10 · confidence 4/52024-07-29

Summary

This paper introduces a new type of CATE estimator using instrumental variables. The proposed method employs the direct learning approach.

Strengths

1. The paper is self-contained and comprehensible. 2. Besides developing the CATE estimator, the paper also proposes an estimator for finding the optimal treatment regimes.

Weaknesses

__Missing literature review__ The paper most closely related to this work is [Machine Learning Estimation of Heterogeneous Treatment Effects with Instruments[(https://proceedings.neurips.cc/paper/2019/file/3b2acfe2e38102074656ed938abf4ac3-Paper.pdf). It develops a fast-converging CATE estimator for local average treatment effects using instrumental variables. However, this paper is not cited in the literature review. Please consider including a discussion of this paper for richer context. More importantly, please compare your work with this paper to highlight the novelty of the current work. __Weak motivation on directed learning__ The introduction section lacks plausible reasons for proposing directed learning. What are the alternative methods and their pros and cons? Why should we specifically consider the directed learning approach? __Validity of Assumption 2__ Assumption 2f is a weaker version of the following assumption: "$U$ is noninformative to $A$ given $Z$ and $X$ (i.e., $A \perp U \mid X,Z$)." Given that there are no practical settings where Assumption 2f holds while $A \perp U \mid X,Z$ doesn't (except some peculiar parametrization), and both assumptions are non-testable, I don't see any practical distinction between $A \perp U \mid X,Z$ and Assumption 2f. In other words, Assumption 2f is just another representation of $A \perp U \mid X,Z$ tailored for identification. Combining $A \perp U \mid X,Z$ with Assumption 2c ($Z \perp U \mid X$) results in $(U \perp A \cup Z \mid X)$ by the contraction property of conditional independence. This means $U$ does not influence $(A,Z)$ given $X$. Consequently, in any related causal graph, there should be no edges from $U$ to $A$. This leads to the ignorability condition that $Y(a) \perp A \mid X$. In summary, interpreting Assumption 2f as $A \perp U \mid X,Z$ means Assumption 2 is essentially an ignitability assumption. Therefore, it is important to discuss the validity of Assumption 2 in practical settings, to disprove that Assumption 2f is merely another representation of $A \perp U \mid X,Z$ designed for identification. Have you considered the LATE setting, given that the estimand in Proposition 1 will remain unchanged? __More analysis is required__ Multiple robustness properties provided in Theorem 3 imply that the proposed estimator converges to the optimal estimator faster. For example, if nuisances converge at an $n^{-1/4}$ rate, where $n$ is the number of samples, then the estimator converges at an $n^{-1/2}$ rate. These results are beneficial since they guarantee fast convergence. While Theorem 3 is attractive, it is somewhat impractical because, in practice, the working model is rarely considered a true model. Please provide more analysis on the rate of convergence concerning the convergence rate of nuisance parameters. __Fair comparison with other estimators__ Even if the empirical evidence in Table 1 is strong, the discussion on why the proposed estimator converges faster than other multiply-robust estimators, such as MRIV, is missing. Asymptotically, there are no reasons to believe the proposed estimator converges faster than the MRIV estimator. Can you provide a discussion on why the proposed estimator converges faster than its competitors?

Questions

1. $\Delta(x)$ in line 89 and $\Delta(x)$ in line 94 are the same? 2. What are the practical examples where Assumption 2 holds? 3. Is there a reason to choose Assumption 2 other than the LATE assumption? Both assumptions yield the same target parameter (in Proposition 1).

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

2

Limitations

1. The paper assumes discrete/binary $Z$. 2. The paper is relying on Assumption 2.

Reviewer M5di2024-08-11

Response

Thank you for your response. My concerns about the motivation and the justification of Assumption 2f have been addressed. However, the following questions remain: 1. Can you provide a discussion on why the proposed estimator converges faster than its competitors? 2. What is the rate of convergence in relation to the convergence rate of nuisance parameters? I believe the paper could be stronger if 1. the motivation of the direct learning is more clearly explained in the introduction section, and 2. the rate of convergence is added. By the way, the current answer misses the response for my question: Can you provide a discussion on why the proposed estimator converges faster than its competitors?

Authorsrebuttal2024-08-12

Response to Comment

Thank you for informing us that your concerns about the motivation and justification of Assumption 2f have been resolved. We value your continued engagement and feedback. Below are our responses to the remaining questions and suggestions: ## Remaining Questions on Convergence Rate We would like to clarify that we did not claim the proposed estimator has a faster convergence rate. Our theoretical results only establish consistency. We will review our text to ensure it does not suggest that we provide convergence rate results. While we are not aware of similar findings in the existing literature, especially related to Wang and Tchetgen Tchetgen (2018), this could be an intriguing direction for future research. ## Suggestions to Strengthen the Paper 1. **Motivation for Direct Learning in Introduction:** We will move the explanation about the benefits of direct learning over traditional methods like Q-learning to the introduction section as suggested to improve the motivation. 2. **Including Rate of Convergence:** We agree. As mentioned, while we are not aware of similar findings in the existing literature, especially related to Wang and Tchetgen Tchetgen (2018), this could be an intriguing direction for future research.

Reviewer M5di2024-08-13

Response

Thank you for addressing my concern. I will raise my point from 5 to 6, given that the end of the discussion period is coming. My final question is this: Can you discuss the fast convergence shown in the experiment? In theory, MR-IV and the proposed estimators both converge fast, but the proposed estimator outperforms in the simulation. Can you justify this simulation result?

Authorsrebuttal2024-08-13

Response to Question

Thank you for your feedback and for raising your score. We understand that you would like some explanation for the better performance of the proposed method for the finite samples in the simulation. One possible reason is that our proposed IV-RDL1 method requires fewer nuisance estimates than MRIV. Another reason may be that in all our proposed methods, we use inverse propensity score $1/\pi_Z(Z,X)$ as the weights in a weighted least square framework. This helps to balance different IV groups and helps to reduce bias and variance (see Seaman and White, 2013).

Reviewer M1zT2024-08-12

One quick question

Thank you for your response. Regarding the binary case, you mentioned, "During our research, we discovered that extending our framework to accommodate other types of treatment is a non-trivial task." Could you please provide more insight or intuition on why this is the case?

Authorsrebuttal2024-08-12

Response to Question

Thank you for your follow-up question. We appreciate your interest in understanding the challenges associated with extending our framework to accommodate other types of treatments. We will briefly highlight the key aspects of identification in the binary case and then explain the challenges involved in generalizing to other types of treatments. To begin with, we will need the following notations: - $\Delta(X)=E[Y(1)-Y(-1)\vert X]$ - $\delta_Y(X)=E[Y\vert Z=1,X]-E[Y\vert Z=-1,X]$ - $\delta_A(X)=P[A\vert Z=1,X]-P[A\vert Z=-1,X]$ - $\tilde\delta_Y(X,U)=E[Y(1)-Y(-1)\vert X, U]$ - $\tilde\delta_A(X,U)=P[A\vert Z=1,X, U]-P[A\vert Z=-1,X, U]$ The proof of Proposition 1 demonstrates that, in the binary case, the identification on the Conditional Average Treatment Effect (CATE), denoted $\Delta(x)$, hinges on the following relationship: $$\delta_Y(X)=E_U[\tilde\delta_Y(X,U)\tilde\delta_A(X,U)]=\Delta(X)\delta_A(X)$$ Here, Assumption 2f provides a sufficient condition for the validity of the second equation. Then we have identified the CATE: $\Delta(X)=\delta_Y(X)/\delta_A(X)$. For a $k$-arm treatment scenario, a natural approach involves selecting one treatment arm as the baseline and defining the CATE as the difference between each of the other arms and this baseline. This results in a CATE vector of dimension $k-1$. Extending Assumption 2f to accommodate this setup and maintain the equality $E_U[\delta_Y(X,U) \delta_A(X,U)] = \Delta(X) \delta(X)$ is not straightforward. In particular, this identification equation will become a system of linear equations, whose solution requires the inversion of a $(k-1)$ dimensional square matrix. Hence, we feel that this would be too complicated to incorporate into the current papers as an additional section; rather, it deserves a separate paper. The challenge increases with the generalization to continuous treatments, as the existing identification relies on differences between conditional means given two discrete IV levels. This suggests the need for novel theoretical frameworks or assumptions tailored to these more complex scenarios.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC