Conditional Outcome Equivalence: A Quantile Alternative to CATE

Conditional quantile treatment effect (CQTE) can provide insight into the effect of a treatment beyond the conditional average treatment effect (CATE). This ability to provide information over multiple quantiles of the response makes CQTE especially valuable in cases where the effect of a treatment is not well-modelled by a location shift, even conditionally on the covariates. Nevertheless, the estimation of CQTE is challenging and often depends upon the smoothness of the individual quantiles as a function of the covariates rather than smoothness of the CQTE itself. This is in stark contrast to CATE where it is possible to obtain high-quality estimates which have less dependency upon the smoothness of the nuisance parameters when the CATE itself is smooth. Moreover, relative smoothness of the CQTE lacks the interpretability of smoothness of the CATE making it less clear whether it is a reasonable assumption to make. We combine the desirable properties of CATE and CQTE by considering a new estimand, the conditional quantile comparator (CQC). The CQC not only retains information about the whole treatment distribution, similar to CQTE, but also having more natural examples of smoothness and is able to leverage simplicity in an auxiliary estimand. We provide finite sample bounds on the error of our estimator, demonstrating its ability to exploit simplicity. We validate our theory in numerical simulations which show that our method produces more accurate estimates than baselines. Finally, we apply our methodology to a study on the effect of employment incentives on earnings across different age groups. We see that our method is able to reveal heterogeneity of the effect across different quantiles.

Paper

Similar papers

Peer review

Reviewer zMNr6/10 · confidence 3/52024-07-09

Summary

The paper introduces a new estimator to obtain conditional quantiles of a treatment effect and argues that it is superior to the pre-existing "conditional quantile treatment effect" (CQTE) estimator. The main difference is that accurate CQTE estimates require accurate conditional quantile models even though the latter can be hard to fit, while the introduced CQC estimator is more robust to the CDF estimators it relies on.

Strengths

Introduces the doubly robust property to estimating conditional quantile treatment effects, a property that many CATE estimators already have. Detailed theory work supporting the claim of robustness of CQC to its component estimators. Exploits existing work on doubly robust CATE estimation by reframing CCDF estimation as a CATE-like problem.

Weaknesses

In 4.1 you test against estimating the CCDFs separately and taking their difference. Do you mean estimating the inverse CDFs / quantile functions and taking their difference i.e. CQTE? If not, what exactly is the procedure here and why not compare against the difference of two quantile estimators. CCDFs are arguably harder to estimate than quantile functions in my experience (they require the additional 'feature' y, have to work with classification losses instead of regression losses, and are more prone to monotonicity violations). So even if the final estimator here is more robust to bad CCDF estimation I'm not convinced it will necessarily be better in practice to the CQTE approach of fitting quantile functions. This is especially true with modern day neural network based quantile models. NW estimators are unlikely to be as good as those on larger / higher dimension datasets, and it's not clear how the theory would hold in cases like that. Isotonization can be risky and create large flat zones for $\hat{g}$. Did you notice any in your experiments? Are there other ways to get monotonicity here, similar to the various strategies to get monotonicity in multi-quantile regression?

Questions

See above

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes

Reviewer CQR96/10 · confidence 4/52024-07-10

Summary

The paper studies an important problem -- treatment effects often have distributional effects and considering conditional mean outcomes is often not the correct objective for evaluating prescriptive interventions. One popular approach to consider heterogeneous effects across a population which allows researchers to determine treatment effects at different quantiles of the study distribution to determine how treatments might effect the worst-off populations. Unfortunately, as the authors note obtaining quantile treatment effects often relies strict smoothness assumptions on the marginal quantile functions and lacks the double robustness properties of CATE estimators. To this end, the authors propose a new estimand to evaluate heterogenous treatment effects that they show too be doubly robust and apply this estimator to both simulational and real-world data.

Strengths

- The paper provides a clean mathematical framework for their estimator and proves useful finite sample properties that should be theoretically convincing for users - Example 1 is a nice simulational example to showcase why typical conditional quantile estimators may provide instable results when the author's CQC estimator is more robust - Proposition 1 provides the double robustness result which has been a useful benefit in finite sample analyses and helps to make the case for using this estimator - Numerical experiments in section 4 help to show case the benefits of this procedure. In particular the second graph in Figure 3 provides good evidence for using this methodology. - Finally, the application to two real-world data sets is helpful to showcase the viability of this methodology for practioners.

Weaknesses

- I understand that the CQTE can be written as a function of $\Delta^*$ but this does not help me to understand how I should interpret this quantity in relation to the typical CQTE. The authors indicate that this it is a "rephrasing" but and provide the mapping. but its not obvious how this translates in actual examples. Since the authors are proposing a non-standard estimand they should provide more details linking this to the standard analysis. One place where this would be very beneficial is in the interpretation of the real-world data sets. I consider this to be the main weakness of the work because without a clear interpretation for the estimand practioners will not adopt this methodology. Without this I doubt the impact of the paper beyound showcasing a technical deficiency in the CQTE estimand due to smoothness requirements. - I found the algorithms a little difficult to understand especially because of the reindexing where you duplicate the data. This seems purely notational and seems like you could simplify the notation here. - The method does not require smoothness of the marginal quantiles, however h must be Holder-smooth. Certainly in many cases smoothness of the marginal may effect smoothness of h. Since the reason for introducing this method is non-smoothness of the marginal quantiles it might be useful to expand on the class of problems where this is an easier problem beyound the example given in Example 1. How does this interact with smoothness of g? - To that end, could you also provide real-world examples where you might expect the marginals to be poorly behaved but the CQC to be well be haved? - Finally, in the the final paper in paper it would be good to have the CQTE estimates in the real world examples so you can draw a contrast between the methods.

Questions

In remark 3 you mention you use a cross-fitting method to improve stability, why is this the case? Duplicating the data yields non-independence? How do i interpret the earnings for groups with the highest/lowest benefits in Figure 4?

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors have addressed some limitations of their work but have not acknowledged possible negative societal impact. From my perspective, work that emphasizes the importance of considering distributional effects is generally societally beneficial. The main limitation that the authors address are primarily on the interplay between the assumption on h and if they could be relaxed to assumptions on g. However, there are several other assumptions and limitations that should be made clear in this work. The authors should potentially consider including limitations in their analysis of the two real-world data sets which might give practitioners a better understanding of how to use this method. Examples of what could be added are given in the weakness section especially with regards to interpretability of results.

Reviewer nAy36/10 · confidence 5/52024-07-11

Summary

The authors propose the Conditional Quantile Comparator (CQC), a function that maps an outcome from the control group in a binary treatment setting to an outcome in the treatment group, such that they represent the same conditional quantiles in their respective distributions. The estimation procedure consists of two stages: first, estimating a conditional contrasting function that examines the difference between conditional CDFs, followed by isotonic regression. The contrasting function is estimated using pseudo-outcome regression, which confers the algorithm with double-robust properties in finite samples. This estimator addresses some of the shortcomings of the Conditional Quantile Treatment Effect (CQTE), which is not robust to errors in the quantiles. The CQC is supported by a theoretical analysis of finite sample convergence rates under smoothness assumptions, as well as by empirical simulations.

Strengths

The authors propose a novel quantile-based treatment effect estimation procedure tailored for skewed outcome distributions. The double-robustness properties in the first stage of the algorithm are desirable as they improve the rate dependence on the nuisance functions ( propensity and local CDF estimates). The paper is clearly written, with sound theoretical results and promising empirical evidence.

Weaknesses

* The motivation for the estimator is somewhat weak. Estimators for CQTEs with double-robustness properties have been already proposed (see [1, 2]). These estimators have a second order dependence on the rate of quantile estimation, as desired. The authors should contextualize their work with respect of exiting literature (which seems to overall be missing in the paper) and compare their algorithm with the existing techniques. * Moreover, the pseudo-outcome estimation technique in [1] is somewhat simpler (see their Appendix B) since they only require one pseudo-outcome regression. There seem to be some tradeoffs, e.g. the CQC requires a two stage procedure that includes a prost processing step (the isotonic regression), as well as for Algorithm 1 to be run several times for each value of $y_1$ considered (because the conditional CDFs need to be estimated at each value). On the other hand, the robust CQTE algorithm from [1] require estimating the conditional PDF at several points. Overall, the CQC seems to be a more computationally (and statistically) intensive estimation procedure and the authors should at least discuss the tradeoffs. * Furthermore, [1] provides rates for more general classes a functions outside of the linear smoothers framework in [3]. I suggest the authors try to provide similar guarantees in order to generalize their results beyond linear smoothers which, to the best of my knowledge, are not often used in practice. This would be particularly useful since the rate results in section 3.4 require all nuisances to be estimated using linear smoothers. * Empirical evaluations should also include these existing methods. Overall, I tend towards soft rejection. This estimator might be useful in its own right, but the authors should contextualize and compare with existing work. I am willing to update my score upon further discussion. [1] Kallus, Nathan, and Miruna Oprescu. "Robust and agnostic learning of conditional distributional treatment effects." International Conference on Artificial Intelligence and Statistics. PMLR, 2023. [2] Leqi, Liu, and Edward H. Kennedy. "Median optimal treatment regimes." arXiv preprint arXiv:2103.01802 (2021). [3] Kennedy, Edward H. "Towards optimal doubly robust estimation of heterogeneous causal effects." Electronic Journal of Statistics 17.2 (2023): 3008-3049.

Questions

See weaknesses section.

Rating

6

Confidence

5

Soundness

3

Presentation

2

Contribution

2

Limitations

The authors have adequately adressed the limitations of their work.

Reviewer bJi96/10 · confidence 2/52024-07-17

Summary

The authors propose a new estimand called the Conditional Quantile Comparator (CQC), which computes the quantiles of the outcome distributions of the treatment effect, hence improving on the (usual) conditional average treatment effect. The CQC is defined as a measurable function that maps an untreated outcome to the equivalent treated outcome in the same quantile, conditional on covariates. The authors introduce a doubly robust estimation procedure for CQC, where the estimation is re-framed as a CATE problem using the pseudo-outcome framework. The authors present finite sample bounds on the estimator's error, and provide numerical simulations and result on a real employment dataset.

Strengths

- The paper is well written and easy to follow - The paper provides a thorough theoretical analysis for the finite sample bounds and proofs of convergence rates for CQC and proposed estimation algorithms - The simulations and real use-case considered provide a good intuition to the reader on the benefits of the proposed approach

Weaknesses

First of all, I'd like to say I am not very familiar with the causal inference literature, so some of my comments might not be as relevant. - Complexity of using quantiles: can the authors comment in terms of error bounds and convergence rates what is the price to pay to estimate quantiles as opposed to CATE? Some of these results, including the use of pseudo-outcomes and isotonic projections, appear a little bit unusual, as usually quantiles are estimated through the pinball loss. As usual what is important is assessing the left tail or right tail (i.e., whether the treatment effect crosses the zero or not), I would expect the right/left tail quantiles to be the most important ones -- is the error in estimating all the quantiles the same across the entire distribution? - Dimensionality of the input $\textbf{x}$: both the simulation and the real dataset use a very low-dimensional $\textbf{x}$, and the comments on using the Nadaraya-Watson estimator and kernel smoothing are really only applicable when X is generally low-dimensional. How does in practice your method behave as function of the input dimension $\textbf{x}$, and would your comments still be applicable even in higher-dimensional settings? - Estimating quantiles vs entire distribution: how does this approach compare to other approaches that estimate the entire conditional density, such as [1]? Is the use of a doubly robust framework guarantee the optimal error rates? [1] Zhou, T., Carson IV, W. E., & Carlson, D. (2022). Estimating potential outcome distributions with collaborating causal networks. Transactions on machine learning research, 2022.

Questions

See the weaknesses section above.

Rating

6

Confidence

2

Soundness

3

Presentation

3

Contribution

2

Limitations

The authors do address some limitations and implications of their approach.

Reviewer bJi92024-08-11

I thank the authors for their comments. In particular, please include the discussion about the complexities of using quantiles in the paper, as it was not clear to me that inverting the CCDF is actually not that bad in terms of an idea in practice and in terms of error rate. Due to the author points in this response and in the global rebuttal, I increase my score within the limits of my understanding and experience with the subject.

Authorsrebuttal2024-08-12

Thank you for the reply. We will make sure to include a further discussion on the overall estimation difficulty of the CCDF when compared to quantile estimation in the final paper.

Reviewer nAy32024-08-12

Thanks for your response. My concernes have been addressed. I have raised my score with the expectation that the authors will include this discussion (especially the comparisons to Kallus et al. (2023)) in their final version.

Authorsrebuttal2024-08-13

Thank you for the reply. We will make sure to include the comparisons with Kallus et al (empirically, methodologically, and theoretically) in the final paper.

Reviewer zMNr2024-08-13

Thanks for the response! Appreciate the extra info and I feel more confident in my already positive score.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC