Exogenous Matching: Learning Good Proposals for Tractable Counterfactual Estimation

We propose an importance sampling method for tractable and efficient estimation of counterfactual expressions in general settings, named Exogenous Matching. By minimizing a common upper bound of counterfactual estimators, we transform the variance minimization problem into a conditional distribution learning problem, enabling its integration with existing conditional distribution modeling approaches. We validate the theoretical results through experiments under various types and settings of Structural Causal Models (SCMs) and demonstrate the outperformance on counterfactual estimation tasks compared to other existing importance sampling methods. We also explore the impact of injecting structural prior knowledge (counterfactual Markov boundaries) on the results. Finally, we apply this method to identifiable proxy SCMs and demonstrate the unbiasedness of the estimates, empirically illustrating the applicability of the method to practical scenarios.

Paper

Similar papers

Peer review

Reviewer rY6N6/10 · confidence 4/52024-07-10

Summary

The manuscript introduces exogenous matching, an importance sampling method for efficient estimation of counterfactual expressions in various settings. This method transforms the variance minimization problem into a conditional distribution learning problem, allowing integration with existing modeling approaches. The authors validate their theoretical findings through experiments with different Structural Causal Models (SCMs), showing competetive performance in a range of counterfactual estimation tasks. They also examine the impact of structural prior knowledge and demonstrate the method's unbiased estimates and practical applicability in identifiable proxy SCMs. Update: revising score upward following author rebuttal.

Strengths

The topic is timely and important, as counterfactual estimation has become an increasingly popular subject in statistics and machine learning. The proposal builds on recent results in neural causal models, specifically with normalizing flows. The ability to incorporate prior knowledge in the form of Markov boundaries is especially welcome, since computing counterfactuals is often intractable without such constraints. The theoretical results appear sound (though I confess I did not go closely through the proofs) and the empirical results are compelling.

Weaknesses

The manuscript is not always clear, probably because a great deal of material has been moved to the appendix to accommodate page count. The result is a somewhat disjointed text that would likely be better served by a journal publication than a conference paper. That said, I am generally supportive of this submission and would be willing to revise my score upward if my questions are adequately addressed (see below).

Questions

When we say the causal model is “not fully specified”, does that just refer to the structural equations or to the graphical structure as well? In general, I was not always certain just how much causal information is used as input to this method. I’m a bit confused by Eq. 5. If this is meant to be a variance estimator, then presumably the RHS should be something like $\mathbb{E}[X^2] – (\mathbb{E}[X])^2$, where $X$ denotes the likelihood ratio $p(u) / q(u)$, correct? This is almost but not quite what we find here. Looking at the appendix, I don’t see why Eq. 48 follows from Eq. 47. Why do we get to drop the square from the first term? What does the constant $c$ denote in Eq. 6? Is it just the entropy of $P$? A brief word on this would help with intuition. Should there definitely be a negative both in front of the expectation *and* within it? I assume the first summand should just be $-\mathbb{E} [ \log Q(u \mid Y_* (u)) ] $? This would look more like the classic cross entropy formula, as indicated in the following line. Eq. 7 also suggests this. On “augmented graphs” – does this “reverse projection” of the ADMG always work? Different DAGs can have the same ADMG, for instance if two latent variables have all the same endogenous children. Perhaps there’s an unstated minimality assumption at work here? What is $m$ in Eq. 13? It is not clear to me from Sect. 5 what the sample size and data dimensionality are for these tasks? In general, this section appears rushed. The performance metrics are also somewhat surprising. If data is simulated, then we presumably have ground truth with respect to counterfactual probabilities. If so, then why not just compute the mean square error of the proposed estimator, perhaps as a function of sample size?

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

Limitations are adequately addressed.

Reviewer 6eu35/10 · confidence 3/52024-07-12

Summary

This paper introduces an importance sampling method for efficient estimation of counterfactual expressions within general settings. It transforms the variance minimization problem into a conditional distribution learning issue, allowing integration with existing modeling approaches. The paper also explores the impact of incorporating structural prior knowledge, i.e. Markov boundaries, and applies the method to identifiable proxy SCMs, proving the unbiasedness of estimates and illustrating the method's practical applicability.

Strengths

1. The paper is well-structured, with many subsections and bullet points summarizing paragraphs. 2. The approach proposed in this paper has clear intuition and is easy to implement.

Weaknesses

1. Contributions are not disentangled well. All three points involve experimental or empirical findings. 2. Some results of the ablation study are abnormal. First, the results show that the approach proposed is not robust. Under setting SIMPSON-NLIN and M, the inclusion of Markov Boundary Mask significantly improves ESP. However, under setting LARGEBD-NLIN and NAPKIN, the Markov Boundary Mask harms the performance, especially with backbone SOSPF. Second, the ESPs under setting LARGEBD-NLIN with backbone SOSPF are almost 0, even when $|s|=1$, which is hard to explain if including Markov Boundary Mask is effective. Third, the variance under setting LARGEBD-NLIN with backbone NICE is extremely large. 3. Insufficient explanation or legend for figures, making it difficult for readers to understand. For example, in Figure 1, $\mathcal{M}$ with a subscripted hammer is not explained. In Figure 3, the legend does not indicate what different colors mean. 4. Hard to follow. Some terminologies need explanation or reference. For example, in line 234, *faithfulness* is not defined.

Questions

I would like to know whether Theorem 2 and 3 are novel. Have the papers cited (like 81, 3, 94, 111) and other papers proposed methods to obtain Markov boundaries? If so, what is the improvement of the method proposed in this paper?

Rating

5

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors adequately discussed the limitations in the paper.

Reviewer aexG5/10 · confidence 3/52024-07-13

Summary

This paper presents Exogenous Matching (EXOM), a new importance sampling method for estimating counterfactual probabilities in Structural Causal Models (SCMs). EXOM transforms variance minimization into a conditional distribution learning problem, providing an upper bound on counterfactual estimator variance as per Theorem 1. It outperforms existing methods across various SCM settings and integrates well with identifiable neural proxy SCMs for practical applications. By incorporating prior knowledge through Markov boundaries, EXOM further enhances performance, demonstrating its potential as an efficient tool for counterfactual estimation in diverse scenarios.

Strengths

1. EXOM provides a tractable and efficient approach for counterfactual estimation in general settings, including scenarios with discrete or continuous exogenous variables and various observations and interventions. This flexibility makes it applicable to a wide range of causal inference problems. 2. The method is built on solid theoretical grounds, with the authors deriving an optimizable variance upper bound for counterfactual estimators. 3. The authors incorporate structural prior knowledge, specifically Markov boundaries, into the neural networks used for parameter optimization. They empirically validate the effectiveness of this approach across various scenarios. 4. EXOM consistently outperforms other importance sampling methods in various SCM settings, as demonstrated by the experimental results. Its compatibility with identifiable neural proxy SCMs further enhances its practical applicability.

Weaknesses

1. Theorem 1 relies on the assumption that the density ratio $q(\mathbf{u}|\mathbf{y}_ {\ast})/q(\mathbf{u}|\mathbf{y}_ {\ast}^\prime)\leq \kappa$ holds for all $\mathbf{u} \in \Omega_{\mathbf{u}}$ and $y_{\ast}$, $y_ {\ast} ^\prime \in \Omega_{\mathbf{Y}_{\ast}}$. This assumption may be overly stringent, as probability measures with infinite support sets might easily violate it. Could the authors elaborate on this assumption and provide examples of distributions that satisfy it? 2. In the Sampling and Optimization section, the distribution of the exogenous variable $\mathbf{U}$ is assumed to be known. However, in practical scenarios, $\mathbf{U}$ is often unknown, necessitating additional efforts to estimate $\mathbf{P_U}$ [A]​. Could the authors provide further clarification on this assumption and discuss potential methods for estimating $\mathbf{P_U}$​? 3. I understand the authors only consider models that provide identifiability results. However, it is encouraged to include neural proxy SCM methods based on VAE and DDPM as experimental baselines. While these may lack identifiability guarantees, comparing against them would further illustrate the superiority of the proposed method in relation to current state-of-the-art techniques. 4. While the method shows good performance on the tested SCMs, it's unclear how well it scales to larger, more complex causal models. The experiments are conducted on relatively small SCMs, and scalability to high-dimensional or densely connected causal graphs isn't thoroughly addressed. [A] Ren, Shaogang, and Xiaoning Qian. "Causal Bayesian Optimization via Exogenous Distribution Learning." *arXiv preprint arXiv:2402.02277* (2024).

Questions

1. In Table 1, the EXOM method with MAP shows significantly better performance on the SIMPSON-NLIN and NAPKIN datasets compared to EXOM with GMM, whereas the performance on the FAIRNESS-XW dataset is similar for both methods. Could the authors provide further explanation for this discrepancy? Why does the GMM approach underperform in these specific cases, and what factors contribute to the similar performance on FAIRNESS-XW? 2. In the ablation study investigating the impact of injecting Markov boundaries, could the authors please include the performance results of EXOM without Markov boundaries? This comparison would further illustrate the benefit of incorporating Markov boundaries.

Rating

5

Confidence

3

Soundness

2

Presentation

2

Contribution

2

Limitations

The authors discuss the limitations of their work in Section 6 and Section D.3. Notably, this work does not have a negative societal impact.

Reviewer 8MDB7/10 · confidence 3/52024-07-17

Summary

Based on the importance sampling methods, the authors propose an exogenous matching approach to estimate counterfactual probability in general settings. They derive the variance upper bound of counterfactual estimators and transform it into the conditional learning problem. They also employ the Markov boundaries information in the inference to improve the learning performances further. Extensive experiments validate the superiority and practicality of their method.

Strengths

- This paper is clearly and well written. - The authors give a theoretical analysis of their estimator, its log-variance upper bound in the general settings, and the counterfactual Markov boundary, and they also perform extensive experiments to demonstrate effectiveness in several cases: with two types of stochastic counterfactual processes, with three categories of fully specified SCMs, etc.

Weaknesses

- This paper makes it clear in lines 140-144 about the assumptions needed for the proposed method. Regarding assumption ii), I think it is not mild, and I am wondering if the proposed method would be sensitive to the specified distribution $P_{\textbf{U}}$ of $\textbf{U}$. The authors might have performed such experiments, but it is not quite clear. - Are the Markov boundaries learned from the observational data via the d-separation, or they are given prior? The authors claimed that such Markov boundaries are structural prior knowledge in line 223, whereas they gave Theorem 3 to demonstrate how to obtain them. - It is suggested to offer the whole procedure or pseudo code of their proposed algorithm somewhere. - In line 239, “augmentied” might be a typo error.

Questions

Please see my questions in Weaknesses.

Rating

7

Confidence

3

Soundness

4

Presentation

4

Contribution

4

Limitations

Not applicable.

Reviewer rY6N2024-08-12

Re: Author rebuttal

Many thanks to the authors for their detailed replies to my comments. I will be revising my score upward in light of these clarifications.

Authorsrebuttal2024-08-13

We thank the reviewer for reading our rebuttal and updating the score accordingly. If there are any further questions, we would be happy to provide additional information.

Reviewer aexG2024-08-13

Thank you to the authors for the detailed responses. Regarding Weakness 1, in addition to the relaxed assumptions, could the authors kindly provide some concrete examples of distributions with infinite support sets that satisfy these assumptions?

Authorsrebuttal2024-08-13

Thank you for the reviewer’s attention to this weakness. In fact, without concentration inequalities or approximations, it is challenging to find concrete examples that fully satisfy this assumption on infinite support sets. Even the density ratios between Gaussian distributions often violate boundedness asymptotically, as illustrated in [1] (Eg. 4.1). The best example we could find is from [3] (Eq. 11), which provides upper and lower bounds between the binomial distribution and the Poisson distribution (with infinite support) under certain conditions. Of course, once concentration inequalities or approximations are employed (as in relaxed formulations 2 and 3), the weakness regarding infinite support becomes less acute, provided we prove or assume that the probability measure of $P_\mathbf{U}$ on the unbounded portion is zero or very close to zero. ---- [3] [Dümbgen, L., Samworth, R., & Wellner, J. (2021). Bounding distributional errors via density ratios.](https://arxiv.org/abs/1905.03009)

Reviewer 6eu32024-08-13

Thanks for the rebuttal by the authors addresses my concerns. I will update my rating accordingly.

Authorsrebuttal2024-08-14

We are glad that the rebuttal addressed the reviewer's concerns, and we appreciate the reviewer for updating the rating accordingly.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC