High Precision Causal Model Evaluation with Conditional Randomization

The gold standard for causal model evaluation involves comparing model predictions with true effects estimated from randomized controlled trials (RCT). However, RCTs are not always feasible or ethical to perform. In contrast, conditionally randomized experiments based on inverse probability weighting (IPW) offer a more realistic approach but may suffer from high estimation variance. To tackle this challenge and enhance causal model evaluation in real-world conditional randomization settings, we introduce a novel low-variance estimator for causal error, dubbed as the pairs estimator. By applying the same IPW estimator to both the model and true experimental effects, our estimator effectively cancels out the variance due to IPW and achieves a smaller asymptotic variance. Empirical studies demonstrate the improved of our estimator, highlighting its potential on achieving near-RCT performance. Our method offers a simple yet powerful solution to evaluate causal inference models in conditional randomization settings without complicated modification of the IPW estimator itself, paving the way for more robust and reliable model assessments.

Paper

Similar papers

Peer review

Reviewer yvJa6/10 · confidence 3/52023-07-06

Summary

The authors formulate and evaluate an approach to solving a non-standard problem: evaluating a causal model M when additional data (not used to construct M) is available from a non-randomized experiment. In particular, the authors focus on comparing IPW estimates from the non-RCT data and from the inferences of the model.

Strengths

The idea is a simple and apparently powerful one: Remove the variability due to IPW by performing IPW on both the actual data and the estimates from the model. This largely removes an apparently extraneous source of variability (IPW itself) and allows direct comparison of the estimates.

Weaknesses

The basic idea of the pairs estimator assumes some basic properties of the IPW estimator. Specifically, if IPW was a terrible estimator whose estimates were unrelated to the data (for example, it always output a single value for ATE: 0.5), then the pairs estimator would always show that the model was essentially perfect, regardless of the model’s estimates or the non-RCT data. I don’t think this scenario likely, but it is possible. This implies that, at least, some diagnostic tests are in order to increase confidence in the output of the pairs estimator. For example, you could introduce noise into the model’s estimates and see if the estimated error increases. The results in Figures 2 and 3 are very good. Indeed, they are *freakishly* good. They are so good that it makes me wonder whether the experiments reported in these figures are really evaluating anything important about how the pairs estimator works in practice. This bears some discussion in the description of the results. The introduction to the paper may be confusing to many readers. When first encountering the term “non-random experiment”, many readers will balk, thinking that randomization is the *sine qua non* of experimentation and seeing “non-random experiment” as a contradiction in terms. The authors could save readers this confusion by introducing the example of explicit non-random assignment (line 38) earlier or by moving the first paragraph of Section 2 (Related Works) to early in the introduction. Another issue may confuse readers: The contrast between an IPW estimate and the “model’s” estimate of treatment effect. For many readers, the goal of analyzing observational data is to get a single estimate of treatment effect, perhaps through IPW. In this scenario, there is no “model” (or, the model is a model of treatment propensity). The authors could substantially improve the paper by explaining one or more practical scenarios in which a researcher has both a model of causal effect and a non-RCT data set that has not been used to create that model. The paper has occasional small grammar errors that detract from the authors’ message. An example is the first sentence of the contributions (with corrections noted in brackets): “We focus on causal model evaluation with non-RCT[s], propose a novel method for low-variance estimation of causal error (Equation 1), and demonstrate its effective[ness] over current approaches by achieving near-RCT performance.” A little more care in editing would improve the paper.

Questions

1. What are examples of practical scenarios in which researchers have a model of treatment effect and then they collect data non-RCT data to evaluate that model? 2. Under what circumstances would the pairs estimator fail to provide estimates with low variance and low bias?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The experiments do not identify cases in which the pairs estimator will fail. In the synthetic experiments whose results are shown in Figure 2 and 3 (which could be used to identify such failure modes), the results for the pairs estimator are freakishly good.

Reviewer t8vT6/10 · confidence 3/52023-07-06

Summary

This paper proposes a new estimator for the causal error, that achieves lower variance than previous approaches. The estimator consists in the difference between a IPTW-like estimator using the causal model and a direct IPTW causal effect estimator. The paper shows that under clear assumptions, this estimator results in lower variance than a naive estimator, which only uses a typical IPTW. The authors further test this estimator empirically on a wide range of setups that both satisfy and potentially violate their assumptions.

Strengths

This paper propose a simple yet effective way to reduce the variance of the causal error estimation. The assumptions under which the main result hold are very clearly stated and discussed. The authors have extensively tested their approach. In particular, they have tested both on simulated data that satisfy their assumptions as well as on data that potentially violate them, to stress test the method.

Weaknesses

The main theoretical comparison of this paper seems to be the naive IPTW estimator. As the authors state in the related works section, other estimators have been proposed in the literature to reduce the variance of the causal effect estimator. How does this estimator compare to those theoretically ? And can you leverage some the improvements of IPTW in your method too ? I think the whole goal of the paper would deserve more clarity. For instance, the overall goal is usually to estimate the true causal effect rather the causal error. Having a good estimator of the causal effect would directly result in low causal error. I believe this would deserve some more motivation / details in the text. Furthemore, I would encourage the author to make the problem setup more clear. In my opinion, it is not fully transparent from the paper if the model is trained on a different set of samples than the ones used for the IPTW estimator. I believe this is the case but this should be state more clearly.

Questions

Same as above: The main theoretical comparison of this paper seems to be the naive IPTW estimator. As the authors state in the related works section, other estimators have been proposed in the literature to reduce the variance of the causal effect estimator. How does this estimator compare to those theoretically ? And can you leverage some the improvements of IPTW in your method too ? I think the whole goal of the paper would deserve more clarity. For instance, the overall goal is usually to estimate the true causal effect rather the causal error. Having a good estimator of the causal effect would directly result in low causal error. I believe this would deserve some more motivation / details in the text. Furthemore, I would encourage the author to make the problem setup more clear. In my opinion, it is not fully transparent from the paper if the model is trained on a different set of samples than the ones used for the IPTW estimator. I believe this is the case but this should be state more clearly.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The limitations and assumptions of this work were carefully adressed.

Authorsrebuttal2023-08-21

Thank you again for your review, and we would like to ascertain if our rebuttal has adequately resolved the concerns you raised. We continue to welcome any supplementary observations or clarification to bolster our work.

Reviewer zDTG5/10 · confidence 3/52023-07-09

Summary

This paper constructs a new estimator for IPW evaluation by comparing the IPW estimator applied on the model-predicted treatments versus the observed treatments. The paper presents a theoretical result that this estimator has lower variance than the naive one and aims to demonstrate this via empirical experiments.

Strengths

1. The structure of the paper and most of the writing are very good. 2. The theoretical result is important. 3. The empirical results seem to show that the estimator is better than the naive one.

Weaknesses

1. The paper could benefit from more clarity in the writing. For example, in line 124, it would be great if there can be some intuition or example on when P(T=1|X) is skewed and what that means (is it overfitting or mis-specified)? Furthermore, it would be great if there was an example of a commonly used model and how it satisfies Assumption A, and why this assumption is not a strong one. 2. While the theoretical result seems important, the supplementary file with the proof is missing, which makes it impossible to review. 3. Similarly, the empirical experiments, while extensive, rely on datasets and details that are supposed to be in the supplementary file, but that file is missing. Therefore, it is unfortunately impossible to review the setting in detail.

Questions

1. Could the authors briefly describe the derivation of the IPW estimator (the unlabeled equation between eq. 4 and 5). I have not seen it in this form before. 2. How do the authors explain the large spike in variance at \sigma_\beta = 10 in Figure 2 for the baseline?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

3 good

Limitations

The main limitation is the lack of the Appendix which contains significant details regarding the contributions of this paper. The authors should also discuss potential limitations of their theoretical assumptions.

Reviewer r45Q7/10 · confidence 3/52023-07-24

Summary

This work aims to evaluate the fidelity of causal models in estimating true treatment effects across different treatments. The golden approach involves comparing treatment effects derived from the target causal model and those obtained from Randomized Controlled Trials (RCT). Practical, time, cost, and ethical constraints often necessitate replacing the RCT estimate with non-RCT methods such as Inverse Probability Weighting (IPW). However, IPW may lead to unbounded variance due to imbalanced propensity scores. To address this, the authors introduce a procedure that applies the IPW estimator to both the model and the actual effects. This aligns the estimated treatment effects, thus offsetting their estimation errors, and results in a lower variance causal error estimate. Under the two stated assumptions, the authors show that the variance of the causal error estimated from their approach, namely pairs estimator, is upper bounded by variance of the the causal error estimated from the naive estimator. In their experiments, the authors compared their approach with the naive estimator, RCT estimator, and existing state-of-the-art variance reduction estimators such as the self-normalized estimator and the LW IPW estimator. They carried out these comparisons on three synthetic datasets, under various non-RCT scenarios, which included different treatment assignment units across sub-populations and varying degrees of propensity score imbalance. The results demonstrated that their approach consistently produced low estimation errors, often on par with those from the RCT estimator. Moreover, they also evaluated the performance of their approach when the existing machine learning-based causal models are used in treatment effect estimation. Still, the pairs estimator yield low estimation errors and yielded results comparable to those from the RCT estimator.

Strengths

1. The author proposes a simple yet powerful procedure for estimating causal error without modifying IPW, which might lead to other complexity, such as parameter tuning, and bias introduction. Their experiments demonstrate their approach has the capability in estimating true causal error faithfully under many existing non-RCT scenarios that are often encountered in practice. 2. The problem statement, formulation, and illustration are clearly stated and well structured, allowing readers to follow easily.

Weaknesses

Despite their approach being supported by the theoretical results and extensive experiments presented, the authors have not provided an appendix. The thorough justification of their theoretical results and experimental details could only be further substantiated with access to this supplementary material.

Questions

1. Regarding Figure 2, can you clarify why the variance of the Linear-modified estimator and Normalized IPW initially increase and subsequently decrease as the imbalance degree increases? Wouldn't these variance reduction methods yield improved results when the degree of imbalance is less pronounced? 2. In assumption A, what is $b_i$? 3. Given that your theoretical findings strongly depend on Assumption A, the compliance of the learned causal model's counterfactual predictions with this assumption becomes a key aspect of your experimental inquiry. Could you please provide the appendix and discuss these results in more detail? #################################################################################### [08/19/ 2023] Reviewer r45Q: The experiment validation on Assumption A and the proof for Proposition 1 are provided, and hence I adjust my review accordingly. ####################################################################################

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

To my knowledge, this work does not have potential negative societal impacts. However, the authors did not provide a section with these discussions.

Reviewer PdWH6/10 · confidence 3/52023-07-30

Summary

The paper considers the estimation of causal error using the IPW estimator in conditional randomized experiments. Given that the allocation probabilities are readily available in these types of experiments, IPW estimators are often used. The authors propose to use the same IP weights for both the causal prediction, and the the estimator of the ground truth, and show that this so-called pair estimation approach reduces the variance in the distance between the causal prediction and ground truth. They provide theoretical justifications for their approach in terms of variance reduction, and illustrate this on synthetic data sets.

Strengths

The idea of using the same IP weights for both the causal prediction and ground truth estimation is simple yet effective.

Weaknesses

There is no evaluation of real data sets. Often the IPW is bad when the ps is unknown and the model is misspecified. When the ps is known by design, this seems to be less of a concern, and can/should probably be addressed by a better design. Often the IP weights are used in conjunction with an OR model to construct a DR estimator. In fact, there is really no good reasons to use the naive IPW estimator considered by the authors. I doubt the proposed method may also be useful for improving the DR estimator, but not as dramatic. The authors should probably have done this.

Questions

The authors refer to the phenomenon that the oracle IPW may not be as efficient as the one using a correctly specified model as one of the limitation of the IPW estimator. Why is this a limitation?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The author’s use of non-RCT data is very misleading; this is actually often called the conditional randomized studies (e.g. Imbens and Rubins’ book), and is one common type of randomized studies. A trial does not need to have the same coin for everyone! The authors mention that their approach is different from the modern approaches that try to stablize the IP weights. It is unclear to me if one already use these stablization methods, then whether the authors’ method is still useful.

Reviewer r45Q2023-08-19

The experiment validation on Assumption A and the proof for Proposition 1 are provided, and hence I adjust my review accordingly.

Authorsrebuttal2023-08-21

Thank you again for your feedback and we appreciate the recognition of our efforts in addressing the concerns raised.

Reviewer zDTG2023-08-20

I appreciate the authors' response, I believe it addresses my questions and concerns. I am adjusting my score accordingly.

Authorsrebuttal2023-08-21

Thank you for your valuable input, and we're grateful for the recognition of our efforts in tackling the issues brought up.

Reviewer yvJa2023-08-20

Thanks for the additional information and thoughtful responses. I am increasing my rating.

Authorsrebuttal2023-08-21

We sincerely thank you for your constructive feedback and appreciate your acknowledgment of our attempts to address the concerns raised.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC