Summary
This paper develops a meta-learner-based approach for conformal inference of individual treatment effects (ITEs). Instead of predicting the missing outcome (the counterfactual), the proposed approach uses plug-in pseudo outcomes as the inference proxy. The validity of ITE inference is based on careful analysis of the statistical dominance of the ITE and the inference proxy. In this way, one no longer needs to adjust for the covariate shifts in applying conformal inference, and the validity holds in a model-free fashion (although the coverage guarantee is more limited). Overall, this paper finds a novel approach to predictive inference of ITEs, and the method works well in synthetic and real-world datasets.
Strengths
1. Novel contribution to an important problem
This paper makes novel contributions, with a new methodology, to an important problem: uncertainty quantification and predictive inference of individual treatment effects. The proposed methodology significantly differs from the PO-prediction-based approaches in the literature, and the analysis contains interesting findings/perspectives. I believe these are important contributions to the literature of conformal inference, and might inspire more developments in the future.
2. Good writing quality and clarity
This paper is well-written and enjoyable to read. The challenges are clearly stated and the contributions are easy to capture. I still have some questions/suggestions for improving the presentation; please see my questions.
Weaknesses
1. The analysis is somewhat limited to the known propensity case.
Since previous ITE methods already achieve validity in the known propensity case, the contribution of this paper is somewhat marginal in the sense that it does not push the limit of how well/ how much we can do for this problem (although i still appreciate the novel perspective provided here). However, the authors only mentioned the challenge of unknown propensity score at the end of the paper, which is a bit unsatisfactory.
2. The setting in synthetic datasets should be stated more clearly.
I think the most challenging part of ITE inference is when both POs are missing, and this is where the proposed approach becomes most interesting. However, when reading the experiments part I am not sure which situation we are in. Are we inferring the ITE for both POs missing case, and what is the calibration data here? I think further clarification will be helpful.
3. The benefits of this method in experiments is a bit unclear.
The experiment part contains rich information. However, I am not sure I fully understand in which cases the proposed method performs the best - is it particularly beneficial for both-PO-missing case, or one-PO-missing part? I thought in the latter case the previous approach should suffice, or it still can be improved due to no covariate shift adjustment? I also do not understand why IPW and DR lead to so long prediction intervals for the NLSM dataset. It is mentioned that "conformity scores have “very strong” dominance over oracle scores", but it is difficult to understand in which cases this might happen (strong signal? large noise? high nonlinearity?).
Questions
1. What if the propensity score is unknown?
A natural question is whether the proposed method is robust to fitted propensities in the pseudo outcomes. I guess the dominance condition should still be (approximately) reasonable if the estimation is accurate (at least for DR?) since the expectation is the same, and the linear relation still holds? Currently the discussion on this point is a bit passive. But given the existing methods' good performance in this aspect, I was wondering whether more can be said about this problem.
2. Do we have a sense of how large $\alpha^*$ is for conditions (ii) and (iii) in Theorem 1, or practically used conformity scores?
It is also discussed that $\alpha^*$ is difficult to know, which I think is a hurdle to the current method. While it's said that $\alpha^*$ is evaluated in the experiments, I didn't understand how this can be done. Can we have some examples of the $\alpha^*$ under some parametric models + gaussian noise? Can we know how it changes with characteristics of the distributions?
3. A clarification question: Does it only apply to the setting without covariate shift?
From my understanding the method requires that the calibration data are exchangeable with the future point. This means for both-PO-missing parts we should also find such individuals in the calibration data, and similarly for single-PO-missing ones. If this is the case, it will be helpful to explicitly mention this because it would be helpful to practitioners.
4. Clarification in the presentation of the experiment part.
When reading the experiments part I am not sure which situation we are in. Are we inferring the ITE for both POs missing case, and what is the calibration data here? Did you evaluate the performance for the case with only one PO missing?
5. What affects the relative performance of meta-learner-based approach?
I see that the performance of the proposed approach is not always the best, and it can yield very long intervals for the NLSM dataset. It is important for practitioners to understand in which cases this method will perform well or poorly. Is it possible to give some guidance on how the performance changes with the data generating process, such as the coupling of POs, the strength of signal/noise?
---
Minor points & typos:
1. Should condition (ii) in Theorem 1 be $V_\varphi \succeq_{(2)} V^*$? Seems the current order is opposite to other types of conditions.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.