Quantifying Aleatoric Uncertainty of the Treatment Effect: A Novel Orthogonal Learner

Estimating causal quantities from observational data is crucial for understanding the safety and effectiveness of medical treatments. However, to make reliable inferences, medical practitioners require not only estimating averaged causal quantities, such as the conditional average treatment effect, but also understanding the randomness of the treatment effect as a random variable. This randomness is referred to as aleatoric uncertainty and is necessary for understanding the probability of benefit from treatment or quantiles of the treatment effect. Yet, the aleatoric uncertainty of the treatment effect has received surprisingly little attention in the causal machine learning community. To fill this gap, we aim to quantify the aleatoric uncertainty of the treatment effect at the covariate-conditional level, namely, the conditional distribution of the treatment effect (CDTE). Unlike average causal quantities, the CDTE is not point identifiable without strong additional assumptions. As a remedy, we employ partial identification to obtain sharp bounds on the CDTE and thereby quantify the aleatoric uncertainty of the treatment effect. We then develop a novel, orthogonal learner for the bounds on the CDTE, which we call AU-learner. We further show that our AU-learner has several strengths in that it satisfies Neyman-orthogonality and, thus, quasi-oracle efficiency. Finally, we propose a fully-parametric deep learning instantiation of our AU-learner.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer 8SMd3/10 · confidence 5/52024-06-18

Summary

The paper provides the orthogonal estimator for the distributional treatment effect (the conditional CDF of $Y[1]-Y[0]$).

Strengths

This paper is technically strong, demonstrating a high level of mathematic rigor and diligence. It presents an in-depth description of the proposed estimator. I think the proposed estimator is useful in practice.

Weaknesses

Despite the strong technical details, the paper is poorly written in overall. __1. Weak motivation__ A current shape of Introduction is weakly motivated. Firstly, the way that the paper motivates the problem is misleading: > Methods for quantifying the aleatoric uncertainty of the treatment effect have gained surprisingly little attention in the causal machine learning community. This sentence is incorrect, since there are literatures on quantile regression and semiparametric density estimation, as reviewed in Section 2. More importantly, the introduction doesn't provide the motivation of the problem against the following question: _why a community need the proposed estimator, given that the quantile estimator can capture the distributional treatment effect_. __2. Difficult to understand due to insufficient information__ Another issue that the paper has (especially in Introduction) is its lack of back-ground information that readers may need to comprehend. Specifically, in Introduction, the paper doesn’t provide any definition or clue what CDTE is. I understand that the CDTE is defined in the caption of Figure 1 as $P(Y[1] - Y[0] \leq \delta \mid x)$. However, $Y[a]$ is undefined, and this key quantity should be in the main body of the text. Also, even if Figure 1 aims to provide a whole summary of the paper, authors at Introduction have insufficient knowledge to comprehend it. In other words, Figure 1 is too detailed to be presented in Introduction section. Since Figure 1 can only be understood by those who entirely digested the paper from the beginning to the end, Introduction section is not the right position where the Figure 1 is located. I understand the goal of Figure 1, but it doesn't achieve the goal because of insufficient background information. The same issue happens to Figure 2 and Table 1. For example, in Table 1, technical terms like AIPTW, hold-out residual, optimization assumptions are undefined, so it's hard to appreciate the contribution of this paper. __3. Fuzzy description on contribution__ First, the terms like "Aleatoric uncertainty" and "distributional treatment effect" are used for denoting the same target quantity. Given that the "distributional treatment effect" is clearly describing the problem, I don't see why the authors want to use "Aleatoric uncertainty" as a title and employ these two words to denote the same estimand. Second, the contribution is wrongly described. Consider this sentence: > AU-learner solves all of the above-mentioned challenges 1 – 3 . The AU-learner doesn't address Challenge 1 and 2. Challenge 1 means that the distributional treatment effect is not identifiable. The Makarov's bound, _NOT_ AU-learner, is employed to address Challenge 1. Challenge 2 is actually the same as Challenge 1, since it means that there are no known nuisances-based functional for the distributional treatment effect. Again, the Makarov's bound is used to address Challenge 2. AU-learners are representing the approximated quantity of the target estimand in terms of nuisance functionals. Finally, third contribution "flexible deep learning instantiation of our AU-learner" is scarcely described only in Section 5. The description needs to be much improved. __4. Little focus on the real contribution__ The real contribution of this paper, compared to the existing works in Table 1, is to provide the doubly robust conditional distributional treatment effect for the bounds of the conditional CDF of treatment effects. However, little focus and efforts have been made for this contribution. For example, if developing a doubly robust estimator is a contribution, then corresponding results such as detailed error analysis, a closed form of estimators, a detailed recipe of the proposed estimator for the specified working model, how to minimize the losses in Equations (8,9), a simple example, and assumptions should be described.

Questions

1. Are rate-doubly-robustness and Neyman orthogonality violated when the scaling hyperpameters are not $1$? 2. Is the Makarov bound sharp? 3. What are the practical examples of the distributional treatment effect? 4. In line 178, is this phenomenon officially termed the selection bias?

Rating

3

Confidence

5

Soundness

4

Presentation

1

Contribution

2

Limitations

The paper is limited to the setting where the ignitability holds.

Authorsrebuttal2024-08-12

[1/3] Thank you for the quick response! We sincerely appreciate the time and effort you have taken to provide us with valuable feedback. We are very sorry for the ambiguities in the terminology across existing research research streams and for our original manuscript not providing enough context to resolve them. We further apologize if our previous comments were unclear or could be interpreted as implying a lack of understanding – this was not what we meant, and **we apologize if our response may have come across in the wrong way**. Rather, we feel grateful for having received such thorough and knowledgeable reviews that both commended the quality of our paper and also made further suggestions to improve our work for the camera-ready version. ### Regarding the unanswered Q3 and W1 > “Using your framework, my questions and concerns (Q3 and W1) were about what practically interesting scenarios "Aleatoric uncertainty of the treatment effect" can capture that other research streams cannot.” > “why a community need the proposed estimator, given that the quantile estimator can capture the distributional treatment effect.” **We apologize for misunderstanding your questions**. In our paper, we work with the aleatoric uncertainty of the treatment effect in the form of the CDF/quantiles of the treatment effect at the covariate-conditional level. More specifically, we aim at inferring the probability that the treatment effect is less than or equal to a certain value ($\delta$), conditional on the covariates. For $\delta=0$, the latter becomes the covariate-conditional probability of the treatment harm (benefit), i.e., $\mathbb{P}(Y[1] \le Y[0] \mid x )$. _Why is the distribution of the treatment effect relevant?_ Here are three examples of how the covariate-conditional probability of the treatment harm (benefit) is useful in practice and how other research streams are unable to provide this information: 1. **Medicine**. In cancer care, for instance, the conditional average treatment effect may suggest whether a treatment is beneficial _on average_ while it can not offer insights into how probable negative outcomes are. For example, consider a patient with a tumor and a drug for which the conditional average treatment effect is larger than zero, suggesting that the treatment has a benefit _on average_ (= the tumor size reduces on average). However, the average treatment effect does not tell us how likely such a reduction is. Given the randomness of the potential outcomes, there could be a chance that the tumor size will increase after the treatment. Hence, medical practitioners are often interested in understanding the probability of treatment benefit or harm [1,2], which is captured by the distribution of the treatment effect. This allows us to answer questions such as: _what is the probability that the treatment effect is larger than zero_? In the above example, this means how probable is a reduction of the tumor after treatment? Crucially, such questions cannot be answered by distributional treatment effects and require knowledge of _the distribution of the treatment effect_, which is the focus of our paper. 2. **Public health**. As another practical example, we refer to our case study analyzing the effectiveness of lockdowns during the COVID-19 pandemic. Here, policy-makers are interested in knowing the probability that the incidence after a strict lockdown will be lower than or equal to the incidence without it (=probability of treatment benefit) (see Appendix I). 3. **Post-approval monitoring of drugs.** Understanding the aleatoric uncertainty of the treatment effect is also relevant when monitoring the efficacy of drugs post-approval. Here, substantial increases in the aleatoric uncertainty serve as an early warning mechanism for when treatments are not working well for certain subgroups of patients or when the pharmacodynamics are not fully understood for all parts of the patient population. In sum, there are many examples – especially in medicine – where the distribution of the treatment effect is relevant for practice. Importantly, the distribution of the treatment effect is necessary in order to understand the probability of treatment benefit (or of treatment harm). Below, we also discuss why the distributional treatment effects are very different from the distribution of the treatment effect, and why only the latter can answer the above questions. **References**: - [1] Bordley, Robert F. "The Hippocratic Oath, effect size, and utility theory." Medical Decision Making 29.3 (2009): 377-379. - [2] Nicholson, Kate M., and Deborah Hellman. "Opioid prescribing and the ethical duty to do no harm." American journal of law & medicine 46.2-3 (2020): 297-310.

Authorsrebuttal2024-08-12

[2/3] ### Regarding the categorization of causal quantities Below, we respond to your original question about the differences in causal quantities and what appears to be your main concern. We are convinced that the problem is easy to fix for the final version of the manuscript. > “However, as stated in the authors' response, this categorization is not mentioned in the paper, and no clues are provided regarding this categorization.” We apologize that we mentioned the categorization into the three streams of literature only as plain text, while, after reading your comment, we realized that we should have made it more explicit (e.g., by adding a formal categorization via a table). In general, three different streams are relevant to our work: - the AU of potential outcomes (lines 95-97), - the distributional treatment effects (lines 98-99), and - the AU of the treatment effect (lines 101-122). In the submitted version of the paper, we only made a clear cut between the identifiable causal quantities (1.)+(2.) vs. non-identifiable (3.), which was the main distinction we aimed to communicate due to reasons of space. Below, we appreciate the opportunity to explain the rationale behind our categorization. **Action:** We will revise our paper and make the above categorization in our related work section more explicit. > "Furthermore, this categorization seems counterintuitive -- does it really make sense to state that "distributional treatment effects" and "distribution of treatment effects" are very different?" Thank you for asking this important question. There are indeed two major differences between these streams: 1. **Interpretation**. _Distributional treatment effects_ represent the differences between different distributional aspects of the potential outcomes [3]. Hence, they can answer questions like “How are 10% of the worst-possible outcomes _with treatment_ different from the worst 10% of the outcomes _without treatment_?”. Here, the two groups (treated and untreated) of the worst 10% contain, in general, **different individuals**. This is problematic in many applications like clinical decision support and drug approval. Here, the aim is not to compare individuals from treated vs. untreated groups (where the groups may differ due to various, unobserved reasons). Instead, the aim is to accurately quantify the treatment response for each individual and allow for quantification of the personalized uncertainty of the treatment effect. The latter is captured in the distribution of the treatment effect, which allows us to answer the question about the CDF/quantiles of the treatment effect. For example, we would aim to answer a question like “What are the worst 10% of values of the treatment effect?”. Here, we focus on the treatment effect **for every single individual**. The latter is more complex because we reason about the difference of two potential outcomes simultaneously. Hence, in natural situations when the potential outcomes are non-deterministic, both (a) the distributional treatment effect and (b) the distribution of the treatment effect will lend to _very different_ interpretations, especially in medical practice. In particular, the distribution of the treatment effect (which we study in our paper) is important in medicine, where it allows quantifying the amount of harm/benefit after the treatment [4]. This may warn doctors about situations where the averaged treatment effects are positive but where the probability of the negative treatment effect is still large. 2. **Inference**. The efficient inference of the distributional treatment effects only requires the estimation of the relevant distributional aspects of the conditional outcomes distributions (e.g., quantiles) and the propensity score [1]. However, in our setting of the bounds on the CDF/quantiles of the treatment effect, we also need to perform sup/inf convolution of the CDF/quantiles of the conditional outcomes distributions. Hence, while the definitions of (a) the distributional treatment effects and (b) the distribution of the treatment effect appear related, their estimation is very different. As you can see above, the distributional treatment effects and the distribution of the treatment effect are related to different questions in practice and help in different situations. **References**: - [3] Kallus, Nathan, and Miruna Oprescu. "Robust and agnostic learning of conditional distributional treatment effects." International Conference on Artificial Intelligence and Statistics. PMLR, 2023. - [4] Nathan Kallus. “What’s the harm? Sharp bounds on the fraction negatively affected by treatment”. In: Advances in Neural Information Processing Systems. 2022.

Authorsrebuttal2024-08-12

[3/3] > "However, the categorization seems to come out of nowhere and requires justification." Thank you for the question. Our justification for the above categorization of causal quantities is based on the following rationale. 1. **The AU of the potential outcomes:** The distribution of the treatment is non-identifiable, and, hence, we wanted to first distinguish our work from causal quantities that are identifiable. The reason is that the latter can be addressed by point identification, while the former (our problem) must be addressed by partial identification or making stronger assumptions. 2. **The distributional treatment effects:** Here, we want to distinguish our work from causal quantities that are contrasts between AUs of both the potential outcomes. Thereby, we aim to spell out clearly that the distributional treatment effects and our distribution of the treatment effect are both very different interpretationally and inferentially. 3. **The distribution of the treatment effect**: Here, we aim to survey works related to our setting, namely, the distribution of the treatment effect. Importantly, the above categorization is neither universal nor final but we used it as an informal guidance to structure the related work in our paper. If there is a better way to categorize the causal quantities, we would be happy to incorporate it into our paper. **Action**: We will spell out the rationale for the categorization in our Related Work section more clearly. Again, we are sorry for misinterpreting your initial question and for the confusion this has caused. We hope that we have addressed all of your concerns and assure you that they can be easily fixed in our revised manuscript. Should you have any further questions, please let us know so -- we would do our best to answer them promptly. Thanks again for reviewing our submission.

Reviewer d6jE7/10 · confidence 4/52024-07-07

Summary

The authors propose a partial identification of quantiles of the individual treatment effect, which are not point-identifiable in general. The authors justifiably argue that characterizing the distribution of individual treatment effects gives a better idea of the aleatoric uncertainty in a causal-inference problem. The proposed method is doubly robust, works with heterogeneous treatment effects, and requires minimal additional assumptions about the data.

Strengths

The problem of aleatoric uncertainty in causal effects is clearly important, the contribution is significant, and the presentation is solid and concise. * Explanations of aleatoric versus empirical uncertainty are clear. * The numerous diagrams are informative. * Algorithm 1 is also easy to understand and succinct.

Weaknesses

* The benefit and/or novelty of the CA-learner is questionable (also reflected in the results Table 2) and I wonder if its exposition is taking up valuable space. It is an insightful point on lines 182--185 that learning Makarov bounds could benefit from inductive biases of lower heterogeneity than the conditional CDFs. However, it is unclear if the CA-learner loss really incorporates that inductive bias and if so, how much. * The conclusion is a bit grand. Arguably, previous papers like those highlighted in Table 1 have proposed robust methods for quantifying a version of aleatoric uncertainty of causal effects.

Questions

My main question has to do with recent related work. In particular, Ji et al. [52] appear to solve a similar problem, and the quick dismissal of that approach in this paper because they "made special optimization assumptions" needs further discussion. It would be helpful to spell out what these assumptions are in a concrete sense. In doing so, the authors could also discuss whether these two approaches have any fundamental commonalities, or if the partial identifications are expected to be materially different. Finally, it would be nice to see [52] appear as a baseline in the empirical evaluations, although I understand this could be difficult to implement, especially since [52] is relatively recent.

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

3

Limitations

N/A

Reviewer Cpq87/10 · confidence 3/52024-07-10

Summary

In this paper, the authors propose a method to quantify the aleatoric uncertainty of the treatment effect. For this, authors estimated Makarov bounds on the CDF and quantiles of the CDTE, and then showed, how one can build a learner, which has properties of Neyman-orthogonality and double robustness. The authors proved the abovementioned theoretical properties of the resulting learner and demonstrated the usefulness of the proposed approach in a series of experiments on synthetic and real-world data.

Strengths

I think the paper has the following strengths: It - Addresses an important problem in the field of treatment effect estimation; - Proposed a theoretically grounded method, that has useful theoretical properties, to quantify the aleatoric uncertainty of the treatment effect; - Published code;

Weaknesses

I don't have critical concerns about the paper, but I believe the following can improve the paper: - I find the paper logically well-structured, but at the same time challenging to read, as it is too concentrated with technical details. I would suggest authors reconsider the narrative, concentrating more on the conceptual ideas in the main part, and moving technical details to the Supplementary; - The field of treatment effect estimation is quite special and narrow (from my point of view) among the machine learning community, and it is worth adding some clarifications to the terms and notation used. For example, what is $Y$? It is just said that it is a continuous outcome, without any intuitions of what it could be. - I feel that the paper missing the discussion of the alternatives of the proposed approach. For example, why specifically Makarov bounds were chosen, but not other possible alternatives? - In Table 2, it is worth explicitly writing that CNF corresponds to the Plug-in learner (as it was called throughout the paper). I would suggest Plug-in-CNF (like written for IPTW-CNF). - In Line 268 there is a minor typo: Should be Eq.13 and Eq. 14.

Questions

- What are the alternatives for Makarov's bounds? - Are such options like conformal predictions or plain confidence intervals somehow useful to estimate AU in the context of treatment effect? - Compared to the vanilla Plug-in estimator (say Plug-in-CNF), what is the additional computational overhead of AU-learner? - In lines 212-216, it is mentioned that CA-learned still has shortcomings (a) (and a new shortcoming (c)). In practice, how severe is the shortcoming (a) (selection bias) compared to AU-learner? Can you provide a toy example?

Rating

7

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

One of the limitations, not mentioned by authors, is that the approach requires (conditional) Normalizing Flows, and hence might not work well in the high dimensional scenario.

Reviewer bvQh7/10 · confidence 2/52024-07-10

Summary

The authors introduce AU-learner, a method to estimate the conditional distribution of treatment effects (CDTE) and hence capture the variability in the treatment effect. They use Makarov bounds for partial identification and use conditional normalizing flows for estimation. Further, they show that AU-learner satisfies Neyman-orthogonality and double robustness.

Strengths

(S1) The paper is well written and the authors provide a good experimental validation of the proposed method. (S2) The authors are the first to propose a doubly robust estimator for the Makarov bounds on the distributional treatment effect.

Weaknesses

(W1) I would have liked a more comprehensive discussion of how the proposed bounds compare against existing ones in terms of sharpness. For instance, it is not clear to me if the bounds you get are sharper than [58] in the binary outcomes setting. (W2) I think it would be useful to compare the proposed methods against the non-doubly robust estimators. In practical settings both outcome and propensity models are likely misspecified, hence the main advantage of doubly robust estimators are the faster theoretical rates. In this setting (with both nuisances misspecified), it is not clear if the doubly robust version of the bounds is always better in practice and some more experiments would be nice to explore this. (Minor) In the extended related work on partial identification and sensitivity models (line 658) you could also mention more recent approaches that incorporate RCT data together with proxy and instrumental variables, e.g. the works of [1] and [2]. [58] Nathan Kallus. “What’s the harm? Sharp bounds on the fraction negatively affected by treatment”. In: Advances in Neural Information Processing Systems. 2022. [1] Falsification of Internal and External Validity in Observational Studies via Conditional Moment Restrictions. Hussain et al. AISTATS 2023. [2] Hidden yet quantifiable: A lower bound for confounding strength using randomized trials. De Bartolomeis et al. AISTATS 2024.

Questions

(Q1) For a binary outcome, could you comment on how your bounds relate to Kallus [58]? Why don't you consider them as a baseline (say for estimating bounds on the quantiles of the treatment effect distribution)?

Rating

7

Confidence

2

Soundness

3

Presentation

2

Contribution

3

Limitations

Yes

Reviewer U9cv7/10 · confidence 3/52024-07-12

Summary

The authors study the distribution over the individual treatment effect for binary treatments, continuous outcomes, and observed, potentially high-dimensional confounders. The authors build up on prior work on Makarov bounds, to develop a new method to lower/upper-bound the conditional CDF and the quantile function of the treatment effect. They develop a doubly robust learner and perform a theoretical and empirical evaluation. In a series of experiments, They compare several instances of their method with a baseline based on kernel density estimation. While the bound-based methods clearly outperform the baseline, they yield rather similar results.

Strengths

The authors clearly state the problem and derive a viable solution. They first argue that learning estimators of the CDFs of the two potential outcomes, and plugging them into the bound computation does not yield optimal bounds. Instead, they then propose to learn the outcome CDFs which directly target the Makarov bounds. Third, they augment this loss with a term accounting for selection bias, which yields a doubly robust learner. Finally, the authors derive an implementation based on neural normalizing flows. The authors present a great set of supplementary material including additional discussion of related work, additional experiments, and implementation details. The material might be sufficient for a journal publication.

Weaknesses

The plugin estimator based on a conditional normalizing flow seems to perform rather well in the experiments. The CA learner w/o bias correction seems to perform as good as the AU learner (slightly worse in Table 2, slightly better in Table 4 in the appendix). The added complexity of the AU learner, in particular, the one-step bias correction which comes with an additional tuning parameter, renders the practical value questionable. It seems that a lot of work went into the AU learner; but it might still be worth focusing the presentation more on the simpler CA learner (e.g., discussing the role of g in the main doc).

Questions

I do not fully understand the role of the working model g \in G in the CA learner. The Makarov bound is a CDF and G may include all possible CDFs; so if G is rich enough, the CA learner may return the same CDF estimate as the plugin estimator. If we restrict G, why does this cause tighter bounds?

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

NeurIPS Paper Checklist is provided; no concerns.

Reviewer bvQh2024-08-07

I thank the authors for their response, the comparison to previous bounds cleared my doubts. I maintain my original score of accept.

Reviewer Cpq82024-08-10

I would like to thank the authors for their detailed response to my review. My concerns and questions were well addressed. I believe that the revised version of the paper, with the incorporated changes, will be easier for readers to follow. As a result, I am raising the score from 6 to 7.

Reviewer 8SMd2024-08-10

Response

Overall, I am embarrassed by this response. The response is written as though the authors' categorization—(1) Aleatoric uncertainty of potential outcomes, (2) Distributional treatment effects, and (3) Aleatoric uncertainty of the treatment effect—has been clearly established in the paper, and my questions/concerns arise from a lack of understanding of this framework (e.g., " the distributional treatment effects are not the target of our paper,", where the term "distribution treatment effect" is just defined by this response, not the paper). However, as stated in the authors' response, this categorization is not mentioned in the paper, and no clues are provided regarding this categorization. Furthermore, this categorization seems counterintuitive -- does it really make sense to state that "distributional treatment effects" and "distribution of treatment effects" are _very different_? I do understand that they are different, and that the distribution of the treatment effect, P(Y[1] - Y[0]), is not pointwise identifiable due to the fundamental problem of causal inference. However, the categorization seems to come out of nowhere and requires justification. I do understand which research streams this work is located in. Using your framework, my questions and concerns (Q3 and W1) were about what practically interesting scenarios "Aleatoric uncertainty of the treatment effect" can capture that other research streams cannot. I don't believe your response has fully addressed this question/concern yet. I adjusted my score based on this response.

Reviewer d6jE2024-08-11

I thank the authors for their helpful response and maintain my current score.

Reviewer U9cv2024-08-11

Thanks for clarifying my question! I read the other reviews and the authors' response. I agree with some feedback asking to further clarifying the setting and how it compares/differs to related settings. I have no concerns about the validity and novelty of the method, and stick to my rating.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC