An Efficient Doubly-Robust Test for the Kernel Treatment Effect

The average treatment effect, which is the difference in expectation of the counterfactuals, is probably the most popular target effect in causal inference with binary treatments. However, treatments may have effects beyond the mean, for instance decreasing or increasing the variance. We propose a new kernel-based test for distributional effects of the treatment. It is, to the best of our knowledge, the first kernel-based, doubly-robust test with provably valid type-I error. Furthermore, our proposed algorithm is computationally efficient, avoiding the use of permutations.

Paper

Similar papers

Peer review

Reviewer DRRe7/10 · confidence 4/52023-07-04

Summary

This paper proposes Augmented Inverse Propensity Weighted cross Kernel Treatment Test (AIPW-xKTE), which is a doubly robust test with provably valid type-I error based on kernel mean embeddings to test for distributional treatment effect. The paper has one result, Theorem 4.1, showing the asymptotic normality of the proposed test statistic, and demonstrates its performance on synthetic and real datasets.

Strengths

The paper has one clear goal, i.e. to provide a testing procedure for distributional treatment effect. As the authors mention, distributional treatment effect has been an important topic of research for a while, and their proposal, AIPW-xKTE, has several advantages over the previous methods, most prominently that it has an analytical asymptotic null distribution, circumventing the need for permutation to get the null distribution, as well as being doubly robust. The paper is very clearly written, setting out its goal in the backdrop of previous works and carrying out that goal with minimal fuss. The paper only offers one result, Theorem 4.1, but I think its conciseness is its value. I haven't gone through every detail of the proof in the appendix but a scan convinced me of its soundness. The paper was a pleasure to read.

Weaknesses

Perhaps one thing to count against this paper is its lack of novelty, in that its proposal is not something completely novel, rather it combines two ideas, namely cross U-statistics and augmented inverse probability weighting. However, it combines them to good effect and conducts thorough analysis of it, both theoretically and empirically, and I do not think this should count heavily against the paper.

Questions

The bottleneck of this procedure would probably be the estimation of kernel conditional mean embeddings, which has $n^3$ complexity? Perhaps it would be worth looking at speeding this up, through approximate kernel ridge regression methods (e.g. Nystrom method proposed in [Grunewalder et al., 2012] or FALKON in [Rudi et al., 2017]).

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

The conclusion section has listed a few of its limitations and interesting possible directions for future research.

Reviewer NqCE8/10 · confidence 5/52023-07-05

Summary

The paper proposes a test of the null hypothesis that a binary treatment has no effect on the the potential outcome distribution. The test combines ideas of kernel mean embedding, double robustness, and cross U-statistics.

Strengths

Originality -The connection between kernel embeddings of effect distributions and the cross U-statistic appears to be new. -The main difference from Shekhar et al. (2022), who combine kernel mean embedding and cross U-statistics, appears to be the connection to double robustness. -The main difference from Fawkes et al (2022), who combine kernel mean embedding and double robustness, appears to be the connection to cross U-statistics. -For the kernel embedding of the potential outcome distribution, Muandet et al. (2021) use an IPW-style estimator while Singh et al. (2020) use a regression-style estimator. Similar to Fawkes et al. (2022), this paper combines IPW and regression estimators into a doubly robust AIPW estimator. Quality - The results are clear and appear to be correct, with some minor comments given below. Clarity - The paper is well written, especially its appendix. Significance - Ultimately, this is a paper that combines building blocks that have been partially combined before. The combination is well executed, and contributes to the literature.

Weaknesses

I will raise the score if these items are addressed. Statistical concepts -The paper advertises efficiency, which when discussing tests, refers to certain statistical properties. However, the efficiency being described is computational by avoiding permutations. The framing should clarify this. -Asymptotic equicontinuity and Glivenko-Cantelli class are not well explained in the main text. The former is well explained in the appendix, so a pointer would suffice. The latter is not; please provide more explanation of what this condition means and why it is reasonable in this context. -Some statements about the asymptotic variance are too strong or poorly worded: on line 19 “the asymptotic variance…” and line 164 “the asymptotic variance…” Comparisons -The references given for average treatment effect are actually for the local average treatment effect on line 21. Please update here and elsewhere. -The doubly robust kernel mean embedding estimator can be viewed as augmenting IPW (Muandet et al. 2021) with regression (Singh et al. 2020) approaches to kernel mean embeddings of potential outcome distributions, just as AIPW augments IPW with regression approaches to treatment effects. It would be worthwhile to point this out. -It would be good to see brief comparisons to Shekhar et al. (2022) and Fawkes et al (2022) following Theorem 4.1. Notation -Sometimes notation is overloaded, which is unnecessary and a bit confusing. For example, mu refers to a kernel mean embedding, a regression, and something else in the appendix. -Another notation issue is that the norm for beta is not defined in Theorem 4.1, and the cross fitting is poorly explained compared to Algorithm 1. -It is not a good notation choice to write k(w,y) when x and y are variables with specific meanings in the paper. -Finally, replace O(100n^2) with O(Bn^2).

Questions

The authors write “We were unable to control the type 1 error of the test presented in Fawkes (2022)…” What does this mean? Why not include Fawkes (2022), the most closely related work, in the simulations? The authors show computational efficiency, but how about statistical efficiency? In inequality (iii) of line 656, shouldn’t there be a 2 on the last term? I had some issues with the proof of Step 3. Shouldn’t the final expression have lambda_1^2 in (13)? This correction would continue throughout the proof. At the bottom of page 27, how does the previous display imply lambda_1>0?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

There are no issues.

Reviewer gJy96/10 · confidence 1/52023-07-06

Summary

The paper focuses on studying Augmented Inverse Probability Weighting (IPW) for distributions instead of means. The outline of the paper is as follows: 1. The authors provide motivation for the problem. 1. They review several tools used to solve the problem, including Maximum Mean Discrepancy, Conditional Mean Embeddings, Kernel Treatment Effect, xMMD, and the asymptotics of AIPW. 1. The main results for AIPW in Hilbert spaces are presented, and the authors discuss practical details of the proposed test. 1. Discussion of the experiments.

Strengths

The paper is technical but clearly written. Authors prove non-trivial (to me) technical results, which convinces me that if I were to utilize AIPW for distributions, I would choose this particular implementation. More broadly, if I needed to test treatment effects on distributions and lacked access to propensity scores, I would opt for this test.

Weaknesses

The main limitation of this paper is that I struggle to think of a practical scenario in which I would have an interest in testing differences in the distribution of treatment effects. While the authors briefly mention that this question arises in various applications, none of those applications are utilized in the experiments. I'm uncertain whether investigating the effect of specialist home visits on cognitive test scores, beyond an increase in the mean, is a particularly relevant question to explore. It might be more appropriate to compare this method with conditional average treatment effect tests. I can imagine situations where there are heterogeneous treatment effects that, on average, cancel each other out but work in opposing directions within two populations. edit: see the list attached by authors.

Questions

I'd ask authors to discuss applications in which this test would be useful.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

1 poor

Limitations

yes

Reviewer sucV7/10 · confidence 4/52023-07-09

Summary

The paper introduces a test for the treatment effect which also takes distributional changes into account. The test strongly builds upon the recent works ( Kim and Ramdas, 2023) and ( Muandet et al. (2021)). The main novelty arises from extending the test in Kim and Ramdas, 2023 to the setting of treatment effect estimation. In comparison with previous tests (Figure 4) in the literature, the test proposed in this paper is computationally more efficient at a moderate price in power.

Strengths

Testing for treatment effects is an important problem. The proposed test is computationally efficient and appears to be practically useful.

Weaknesses

1) Figure 2 is misleading to some degree. BART and Causal Forests are estimating the mean of the treatment effect and therefore necessarily fail in scenarios III and IV. 2) I am expecting a comparison against KTE ( Muandet et al. (2021)) and the AIPW extension from Fawkes et al. (2022) in the main text. This should come instead of the current Figure 2. As far as I understand from Figure 4, the main contribution of this paper is not a better test in terms of power but in terms of computational run-time. This does not sufficiently come across in the main text. 3) The contribution of the paper feels fairly limited and more like a straightforward extension of Kim and Ramdas, 2023. The theoretical results appear to me to follow from standard well-established arguments. Can the authors comment more on the technical challenges behind the proofs/ similarities to prior works? 4) While 3) itself is not a reason for a low grade, based on this limitation I would expect a more exhaustive experimental analysis of the method demonstrating the practical usefulness and its limitations on real-world data sets.

Questions

- How does the corresponding plot for Scenario (I) look for the setting in Figure 2 and 4? - I am willing to increase my score if the authors can address the shortcomings above.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

See Weaknesses.

Reviewer yX9h4/10 · confidence 4/52023-07-14

Summary

The paper introduces a statistical test to determine whether the distributions of the two counterfactuals are the same. This goes beyond the well-known average treatment effect, which only tries to understand whether the means of the distributions are the same. The first work in this direction was by (Muandet et al., 2021), who introduced the Kernel Treatment Effect (KTE). However, their test statistic is degenerate, which means that they cannot use the CLT to derive an asymptotic threshold, and they need to resort to a permutations approach in order to compute the threshold. This paper introduces the AIPW-xKTE test, which generalizes KTE by including a plug-in estimator in analogy with AIPW vs IPW, and most importantly by using the same approach as the Cross MMD test from (Kim and Ramdas, 2023), which does yield a statistic with asymptotically normal distribution.

Strengths

The contribution of the paper is clear, and the authors do a reasonable job at placing it within the literature.

Weaknesses

The main weakness is that the contribution is not highly novel, in that the test proposed is basically a combination of the KTE test from (Muandet et al., 2021) and the Cross MMD technique from (Kim and Ramdas, 2023). The explanation would be more transparent if some important concepts were clearly defined. See more details in the questions section. There are some more detailed weaknesses that I also point out in the questions section.

Questions

- What is double robustness? It is a relevant concept in the paper, as it is mentioned twice in the contributions part of the introduction, and many more times later on. However, it is never defined. - Line 23: The plug in estimator is not well defined: what does it actually look like? In lines 222-225, the estimator appears again under a different notation (\hat{\beta} instead of \hat{\theta}). The authors currently say that “At this time, not so many choices exist for estimators…“ and they cite a work on this. It would be good to give a more detailed explanation of how the estimator \hat{\beta} is computed. Similarly, it would be helpful to give more insight on how \hat{\pi} is computed. - Theorem 3.1: What is \hat{\pi}? What is \hat{\psi}_{DR}? These quantities have not been defined before as far as I can tell. - Figure 1: Show error bars in subfigure c, to show that the discrepancy from 0.05 can be attributed to a statistical error. The key question that needs to be answered here is: how large does n need to be for the CLT to kick in and for the Type I error guarantee to hold. Without error bars, the current figure does not provide an answer to this question. - Figure 2: Although not as critical, error bars in this figure would be appreciated too. - Table 1: Show standard error. - Comparison with (Muandet et al., 2021): The authors do not show any experimental comparison with the KTE test proposed by (Muandet et al., 2021). They argue that “Due to the fact that the KTE (Muandet et al., 2021) may not be used in the observational setting, where the propensity scores are not known, there is no natural benchmark for the proposed test.” Since AIPW-xKTE makes use of estimators of propensity scores, it seems natural to me to compare it with KTE using the same propensity score estimators. Is there a reason not to do this?

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

2 fair

Contribution

2 fair

Limitations

No limitations

Reviewer sucV2023-08-10

Response to the rebuttal

I would like to thank the authors for their response. I have increased my score to a 7. Furthermore, I agree with the following concern raised by a different reviewer: >The main limitation of this paper is that I struggle to think of a practical scenario in which I would have an interest in testing differences in the distribution of treatment effects. While the authors briefly mention that this question arises in various applications, none of those applications are utilized in the experiments. I'm uncertain whether investigating the effect of specialist home visits on cognitive test scores, beyond an increase in the mean, is a particularly relevant question to explore. It would be beneficial to provide a concrete application as an example to highlight the necessity of this test. However, considering the limited access to public datasets, this might be an ambitious request.

Authorsrebuttal2023-08-18

Following the concerns raised by Reviewer gJy9, we have provided further examples of use of our test in the respective official comment. We would like to thank the reviewer again for their insightful comments.

Reviewer DRRe2023-08-13

Thank you

I have no more comments to make, and would like to keep my evaluation of the paper. Good luck!

Authorsrebuttal2023-08-18

We would like to thank the reviewer again for their time and comments.

Reviewer gJy92023-08-17

Respectfully, I don't think authors have provided a practical scenario or hypothetical in which I would like to use this test. Perhaps this theoretical construction will find its uses in the future. If the only criterion would be technical novelty I would accept this paper.

Authorsrebuttal2023-08-18

We would like to thank the reviewer again for their time. Although we focused on the specialist home visit example as a novel use of our test, we would like to highlight other potential uses, such that - Determining subgroups of patients that respond differently to medication and establish treatment policies, referring the reader to [1]. - Conducting feature selection for discovering treatment effect modifiers, referring the reader to [2] and [3]. - Studying the effect of various features on Google advertisers' spending, which has very heavy tails since there are a few advertisers who spend a lot [4]. For this reason means are not very useful summaries and instead distributional effects make a lot more sense. We are happy to include any of these references if the reviewer considers that they shed light on the usefulness of our test. Furthermore, we can point the reader to [5], whose introduction discusses some specific motivation for studying distributional effects beyond the mean, as well as the vast literature on quantile treatment effects, which treat distributional effects with the exact same motivation while considering a different distributional target. [1] Chikahara, Yoichi, Makoto Yamada, and Hisashi Kashima. "Feature selection for discovering distributional treatment effect modifiers." Uncertainty in Artificial Intelligence. PMLR, 2022. [2] Bellot, Alexis, and Mihaela van der Schaar. "A kernel two-sample test with selection bias." Uncertainty in Artificial Intelligence. PMLR, 2021. [3] Biesecker, Leslie G. "Hypothesis-generating research and predictive medicine." Genome research 23.7 (2013): 1051-1053. [4] Díaz, Iván. "Efficient estimation of quantiles in missing data models." Journal of Statistical Planning and Inference 190 (2017): 39-51. [5] Kennedy, Edward H., Sivaraman Balakrishnan, and Larry Wasserman. "Semiparametric counterfactual density estimation." arXiv preprint arXiv:2102.12034 (2021).

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC