Conformal Meta-learners for Predictive Inference of Individual Treatment Effects

We investigate the problem of machine learning-based (ML) predictive inference on individual treatment effects (ITEs). Previous work has focused primarily on developing ML-based meta-learners that can provide point estimates of the conditional average treatment effect (CATE); these are model-agnostic approaches for combining intermediate nuisance estimates to produce estimates of CATE. In this paper, we develop conformal meta-learners, a general framework for issuing predictive intervals for ITEs by applying the standard conformal prediction (CP) procedure on top of CATE meta-learners. We focus on a broad class of meta-learners based on two-stage pseudo-outcome regression and develop a stochastic ordering framework to study their validity. We show that inference with conformal meta-learners is marginally valid if their (pseudo outcome) conformity scores stochastically dominate oracle conformity scores evaluated on the unobserved ITEs. Additionally, we prove that commonly used CATE meta-learners, such as the doubly-robust learner, satisfy a model- and distribution-free stochastic (or convex) dominance condition, making their conformal inferences valid for practically-relevant levels of target coverage. Whereas existing procedures conduct inference on nuisance parameters (i.e., potential outcomes) via weighted CP, conformal meta-learners enable direct inference on the target parameter (ITE). Numerical experiments show that conformal meta-learners provide valid intervals with competitive efficiency while retaining the favorable point estimation properties of CATE meta-learners.

Paper

References (51)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer SDYU6/10 · confidence 4/52023-07-04

Summary

This paper develops a meta-learner-based approach for conformal inference of individual treatment effects (ITEs). Instead of predicting the missing outcome (the counterfactual), the proposed approach uses plug-in pseudo outcomes as the inference proxy. The validity of ITE inference is based on careful analysis of the statistical dominance of the ITE and the inference proxy. In this way, one no longer needs to adjust for the covariate shifts in applying conformal inference, and the validity holds in a model-free fashion (although the coverage guarantee is more limited). Overall, this paper finds a novel approach to predictive inference of ITEs, and the method works well in synthetic and real-world datasets.

Strengths

1. Novel contribution to an important problem This paper makes novel contributions, with a new methodology, to an important problem: uncertainty quantification and predictive inference of individual treatment effects. The proposed methodology significantly differs from the PO-prediction-based approaches in the literature, and the analysis contains interesting findings/perspectives. I believe these are important contributions to the literature of conformal inference, and might inspire more developments in the future. 2. Good writing quality and clarity This paper is well-written and enjoyable to read. The challenges are clearly stated and the contributions are easy to capture. I still have some questions/suggestions for improving the presentation; please see my questions.

Weaknesses

1. The analysis is somewhat limited to the known propensity case. Since previous ITE methods already achieve validity in the known propensity case, the contribution of this paper is somewhat marginal in the sense that it does not push the limit of how well/ how much we can do for this problem (although i still appreciate the novel perspective provided here). However, the authors only mentioned the challenge of unknown propensity score at the end of the paper, which is a bit unsatisfactory. 2. The setting in synthetic datasets should be stated more clearly. I think the most challenging part of ITE inference is when both POs are missing, and this is where the proposed approach becomes most interesting. However, when reading the experiments part I am not sure which situation we are in. Are we inferring the ITE for both POs missing case, and what is the calibration data here? I think further clarification will be helpful. 3. The benefits of this method in experiments is a bit unclear. The experiment part contains rich information. However, I am not sure I fully understand in which cases the proposed method performs the best - is it particularly beneficial for both-PO-missing case, or one-PO-missing part? I thought in the latter case the previous approach should suffice, or it still can be improved due to no covariate shift adjustment? I also do not understand why IPW and DR lead to so long prediction intervals for the NLSM dataset. It is mentioned that "conformity scores have “very strong” dominance over oracle scores", but it is difficult to understand in which cases this might happen (strong signal? large noise? high nonlinearity?).

Questions

1. What if the propensity score is unknown? A natural question is whether the proposed method is robust to fitted propensities in the pseudo outcomes. I guess the dominance condition should still be (approximately) reasonable if the estimation is accurate (at least for DR?) since the expectation is the same, and the linear relation still holds? Currently the discussion on this point is a bit passive. But given the existing methods' good performance in this aspect, I was wondering whether more can be said about this problem. 2. Do we have a sense of how large $\alpha^*$ is for conditions (ii) and (iii) in Theorem 1, or practically used conformity scores? It is also discussed that $\alpha^*$ is difficult to know, which I think is a hurdle to the current method. While it's said that $\alpha^*$ is evaluated in the experiments, I didn't understand how this can be done. Can we have some examples of the $\alpha^*$ under some parametric models + gaussian noise? Can we know how it changes with characteristics of the distributions? 3. A clarification question: Does it only apply to the setting without covariate shift? From my understanding the method requires that the calibration data are exchangeable with the future point. This means for both-PO-missing parts we should also find such individuals in the calibration data, and similarly for single-PO-missing ones. If this is the case, it will be helpful to explicitly mention this because it would be helpful to practitioners. 4. Clarification in the presentation of the experiment part. When reading the experiments part I am not sure which situation we are in. Are we inferring the ITE for both POs missing case, and what is the calibration data here? Did you evaluate the performance for the case with only one PO missing? 5. What affects the relative performance of meta-learner-based approach? I see that the performance of the proposed approach is not always the best, and it can yield very long intervals for the NLSM dataset. It is important for practitioners to understand in which cases this method will perform well or poorly. Is it possible to give some guidance on how the performance changes with the data generating process, such as the coupling of POs, the strength of signal/noise? --- Minor points & typos: 1. Should condition (ii) in Theorem 1 be $V_\varphi \succeq_{(2)} V^*$? Seems the current order is opposite to other types of conditions.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Yes.

Reviewer V6dY7/10 · confidence 4/52023-07-07

Summary

The paper introduces a framework for inferring treatment effects, merging concepts of conformal prediction and meta-learner. This framework is characterized by its distribution-free validity and coverage guarantees. The authors utilize stochastic ordering techniques to substantiate the framework's validity.

Strengths

The paper is well written and sounded. The innovative approach of employing stochastic ordering to validate the use of pseudo outcomes in conformal prediction is particularly noteworthy.

Weaknesses

The concept of employing stochastic ordering is innovative; however, its practical application is impeded by the fact that the assumptions are inherently uncheckable In comparison to the method proposed in [1], the current method exhibits better performance only in terms of the RMSE of the CATE. However, when considering coverage and average length, it does not surpass the performance of the inexact method. [1] Lihua Lei and Emmanuel J Candès. Conformal inference of counterfactuals and individual treatment effects. Journal of the Royal Statistical Society Series B: Statistical Methodology, 83(5):911–938, 2021.

Questions

(1). Assuming the propensity score is unknown, could we estimate it first and then follow the same procedure outlined in the paper? (2). For the experimental results, it seems your method can be overly conservative. (3). Figure 4(c) could be adjusted to resolve the issue of elements overlapping.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

See weakness

Reviewer vDHA7/10 · confidence 4/52023-07-07

Summary

This paper adresses individual treatment effect (ITE) inference using conditional average treatment effect (CATE) meta learners. The authors propose to wrap the meta-learners with a conformal prediction step so as to obtain valid confidence intervals for the estimated treatment effect. Since the CATE estimator induces a distribution shift wrt conformity scores, the coverage guarantee does not immediately translate into the desired coverage of actual ITE. In a who can do more can do less spirit, the authors point out that ITE coverage is ensured if conformity scores obtained from pseudo-outcomes stochastically dominate the scores obtained from actual « oracle » outcomes. The authors succeed at providing stochastic dominance results for several stochastic order / meta-learner pairs.

Strengths

- the paper is very clear and pleasant to read - the tackled problem is challenging and impactful - the proposed solution is backup by both theoretical and experimental results

Weaknesses

- the method is by-design over-conservative

Questions

In my humble opinion, this is a very good paper. The problem is well formulated, the proposed solution is clearly exposed and easy to reproduce. The solution is theoretically proved to work (under some clearly stated conditions) which is confirmed by numerical studies. Furthermore, the authors have correctly self-identified the limitations of their approach (knowledge of propensity score and hard to assess level range). My only remark is that I would have like to see a synthetic experiment how far is the proposed solution to the optimum/oralce especially in terms of interval length.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

- knowledge of the propensity score is required - validity range of level $\alpha$ not always easy to determine - over-conservatism of the prediction intervals The two first points are already discussed by the authors in the submission.

Reviewer 855y7/10 · confidence 3/52023-07-31

Summary

The paper deals with confidence intervals for the ITE prediction task, where the authors propose to use the framework of conformal prediction for meta learners. The learning of nuisance models for generating pseudo outcomes is done one separate data split, and the final regression models trained to predict pseudo outcomes are evaluated on a validation set to compute the empirical calibration distribution. The proposed approach for calibration removes the issues of covariate shift and inductive biases on nuisance models with the prior approaches of weighted conformal prediction. The authors also provide guarantees on when the confidence intervals for ITE with pseudo outcomes cover the confidence interval for ITE with true outcomes; and further validate their findings with synthetic and semi-synthetic datasets.

Strengths

* ITE prediction is an extremely relevant problem in causal inference as population level effects (ATE, CATE) would not necessarily transfer to the individual level effects. Further, having an estimate of uncertainity along with ITE prediction is crucial for high-stakes application scenarios. Hence, the approach of conformal meta learners is highly relevant for the current challenges in causal inference. * The proposed combination of conformal prediction with meta learners is novel to the best of my knowledge. Further, the ITE coverage results with pseudo outcomes for standard meta learners (DR Learner) are quite significant as it is non-trivial to understand when the predicted confidence intervals with observed information can represent the confidence intervals with counterfactual information. * The proposed conformal meta learner approach is technically sound with theoretical proofs on their validity and empirical justification on a decently diverse set of benchmarks.

Weaknesses

I do not think the paper has any serious weaknesses, but I have listed some of them in the questions section ahead. One major suggestion is that the paper could be written with better clarification for the section on conformal meta learner being robust to covariate shift and inductive biases. The arguments are stated informally and it is a bit hard to follow.

Questions

* I do not understand the argument of authors that conformal prediction with meta learners is not affected by the choice of inductive biases of nuisance models. The pseudo outcomes are a function of the nuisance models, hence the predicted effect (ITE), which is learnt using pesudo outcomes, indeed depends indirectly on the nuisance models. Hence, the choice of nuisance models would affect the empirical calibration distribution, consequently the confidence intervals for ITE. * I do not follow the argument of authors that conformal meta learners are not affected by the covariate shifts. The covariate shift between the control and treatment population should still affect the computation of pseudo outcomes with finite samples; hence the confidence intervals. * Please clarify the assumption of knowing the true propensity model. What challenges would the authors in their analysis if they do not make this assumption.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Yes, the authors have addressed any potential negative societal impact of their work.

Reviewer SDYU2023-08-15

Thank you for your comments. While the answers to my first and last questions are not completely addressing them, I feel they are fair responses. I'll keep my scores as this paper makes valuable contributions yet still some assumptions/conditions remain not fully clear. One last question just out of curiosity: does this meta-learner idea apply to the confounded setting such as in Jin et al. 2023?

Authorsrebuttal2023-08-15

Thanks for your response

Thanks for your response. When hidden confounding exists, the causal effects are not identifiable so the coverage guarantees do not apply. However, one can operationalize the stochastic ordering framework to conduct sensitivity analysis for the meta-learner in a manner similar to Jin et al. 2023.

Reviewer 855y2023-08-16

Good rebuttal!

Thanks for the good response during the rebuttal! My concerns are addressed and I have increased my score accordingly.

Area Chair mT3T2023-08-18

Reviewers, please respond to author's rebuttal.

As a minimum, please acknowledge that you have read the rebuttal and whether it helps to change your rating, as the authors have tried to respond to your comments in the review. Thank you.

Reviewer V6dY2023-08-19

I acknowledge that I have read the rebuttal and I have increased the rating.

Program Chairsdecision2023-09-21

Decision

Accept (oral)

© 2026 NYSGPT2525 LLC