Summary
The authors propose an approach for conditional average treatment effect estimation using instrumental variables, extending existing work in IVA to settings with weak instruments (i.e., low treatment compliance in some population subgroups). In particular, the approach leverages a two-stage estimation setup: first, a biased CATE estimate is computed from the observational data, then corrected via reweighting based on compliance to obtain a final estimate.
Overall, the paper provides good exposition and intuition for a novel approach, and backs it up with convincing theoretical analysis. The theoretical results are fairly insightful, but can be improved with a little more clarification (details below). The empirical results are not quite as well-motivated, and could also be improved with some clarification of the motivation and differentiation from past work. In particular, I felt that better justification of why these datasets/evaluations are the correct ones to test the proposed approach are needed for me to appreciate the paper’s contributions. Furthermore, I’m a little bit unsure if there are sufficient comparisons to baselines in inference under weak IV in the paper as-is (citations below).
Strengths
1. The exposition is clearly written and provides good intuition on IVA (though I am writing from the perspective of someone that has working knowledge of IVA already).
2. The proposed approach is intuitive and well-motivated, with good theoretical properties. Weak instruments are an inevitable problem in instrumental variable methods, so proposals to reduce their downstream impact are a salient area of research.
3. The paper itself is quite clearly written and mostly easy to understand.
Weaknesses
I'd love to see the following points addressed to clear up any misunderstandings on my part:
1. My biggest criticism is about the motivation and coverage of the empirical results. Coverage-wise, while I appreciate the comparison with a vanilla LATE estimator ($\tau^E$ if I followed correctly?), additional comparisons to baselines in learning with weak instruments would help strengthen the paper, such as the ones cited by the authors as closely related [1, 2] — it’s not clear to me the empirical lift provided by the proposed approach compared to these methods. For the rebuttal phase, in lieu of new results, perhaps a precise explanation of how the authors’ proposed approach shares similarities and differs with [1, 2] would be most helpful.
2. The analysis of the 401k dataset results is slightly imprecise — I’m not sure I buy that $\hat{\tau}(x)$ closely tracks with $\tau^E$; I don’t have a prior for what’s “close enough.” Could the authors clarify more precisely (1) the motivation behind experiments on the 401k dataset (beyond extension to real-world data), (2) and why the results show proof-of-concept for the proposed approach?
3. I have a couple of overarching concerns about the theoretical results. As a first-order comment, how do the generalization bounds of the proposed approach compare to similar bounds for IVA/what are the trade-offs compared to past bounds? As a second-order comment, if I know that $\tau^O$ will be biased (i.e., due to unobserved confounding), why is it imperative for estimation error w.r.t. $\tau^O$ to be low as well (maybe this is so that, given a good estimate of $\theta$ — we get a good estimate of $b(x)$, and therefore correct for the bias)?
Nits (points that would improve the paper in my opinion, but are not urgent):
1. There are a few critical derivations where the clarity could be improved somewhat. It took me quite some time to follow how Eq. 4 was derived — while it’s probably okay to reserve most details to the Appendix, some intuition about which terms are being substituted where would be helpful, and making it clear in Eq. 4 that you’re taking the squared difference (example-wise) of the pseudo-outcome of Eq. 3 (as fitted on the intention-to-treat dataset) and a reweighted version of $b(x) + \tau^O(x)$ as fitted on the observational dataset (since it is equal to $\tau(x)$ by definition).
2. I notice that $\hat{\tau}^E$ and $\hat{\tau}^O$ are fitted on different subsets of the data — somewhat reminiscent of cross-fitting based estimators (e.g., [3]). Out of curiosity, is it necessary to fit the two estimators on different data splits for the theoretical guarantees to hold?
[1] Abadie, A., Gu, J., & Shen, S. (2024). Instrumental variable estimation with first-stage heterogeneity. Journal of Econometrics, 240(2), 105425. \
[2] Coussens, S., & Spiess, J. (2021). Improving inference from simple instruments through compliance estimation. arXiv preprint arXiv:2108.03726. \
[3] Kennedy, E. H. (2023). Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics, 17(2), 3008-3049.
Questions
I think my largest concerns were expressed above in the Weaknesses. I’d be happy to raise my score with a thorough and precise response to especially 1-2. Addressing the following might help strengthen the paper even more, but I don’t consider them as high-priority.
1. I’m slightly confused about the interpretation of Lemma 1 — to me, it seems like we have simply chosen to use IPW to estimate the numerator in Eq. 3. Is this understanding correct? If so, why not consider alternative estimators for the numerator (e.g., doubly-robust methods such as [1])? This is not a huge issue — just making sure I parsed the equation correctly.
2. Did the authors consider/evaluate the sensitivity of the proposed approach to different instantiations of the base learners (e.g., something besides RF/T-learner)? This is not a dealbreaker but would be a nice result to have.
3. Re: “The weighting scheme in Equation 4 creates a weighted distribution…” (L172) — what is this a weighted distribution of? The paper goes on to claim that the difference between the weighted distribution and the target distribution creates a transfer learning problem, but I’m having a bit of trouble understanding what these distributions are. Are these distributions defined over the covariates?
4. Re: Assumption 3-4 (realizability of $b(x)$) — could the authors expand on why this assumption is a reasonable one to make, or point to some works that have made similar assumptions? I think I’ve seen versions of the other assumptions, so I buy those, but I’m not sure if assuming $b(x)$ is linear in the representation $\phi$ is too limiting. Similarly, in Sec. 4.2, there’s an assumption that $\tau^O$ and $\tau$ — one of which is biased — have some shared representation. Is there a more concrete reason for why this is reasonable?
[1] Kennedy, E. H. (2023). Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics, 17(2), 3008-3049.
Limitations
The authors address their limitations in the Appendix. As a tiny suggestion, I think it is important to note that the lack of unobserved confounding is statistically unverifiable yet necessary for many causal inference approaches. Overall, I find the discussion of the limitations to be complete and well-written. As a very, very minor nitpick, I do believe it important to acknowledge the limitations of an approach more up-front (often in the conclusion) — it does not detract from my appreciation of the method and relegating discussion of limitations to the Appendix does not feel quite right to me.