Assessment of the conditional exchangeability assumption in causal machine learning models: a simulation study

Observational studies developing causal machine learning (ML) models for prediction of individualized treatment effects (ITEs) seldom conduct empirical evaluations to assess the conditional exchangeability assumption. We conducted a simulation study to illustrate the performance of causal forest and X-learner models under conditional exchangeability violations, in the presence and absence of true heterogeneity, and to assess the utility of negative control outcomes (NCOs) as a diagnostic. We simulated data to reflect real-world scenarios with differing levels of confounding, sample size, and NCO confounding structures, then estimated and compared ITE-defined strata treatment effects on the primary outcome and NCOs across settings with and without unmeasured confounding. When conditional exchangeability was violated, models failed to recover true treatment effect heterogeneity and, in some cases, falsely indicated heterogeneity when there was none. NCOs successfully identified strata affected by unmeasured confounding. Even when NCOs did not perfectly satisfy ideal assumptions, they flagged potential bias in estimates across ITE-defined strata, though not always pinpointing the stratum with the largest confounding. Violations of conditional exchangeability substantially limit the validity of ITE estimates in routinely collected observational data. NCOs serve as a useful empirical diagnostic tool for detecting unmeasured confounding that varies across subpopulations, helping to support the credibility of individualized inference.

Paper

Similar papers

© 2026 NYSGPT2525 LLC