Response to Reviewer eqxp
We appreciate your efforts in reviewing our paper and constructive comments. We have revised our manuscript based on your advice. We address your concerns in the following.
**Q1.** The paper’s rationale for the calculation of the influence function appears to be lacking, as it directly selects a specific submodel. It might be more appropriate to refer to the function derived in this paper as the efficient influence function, given that the previous doubly robust estimator for binary treatment is efficient.
**A.** While the DR estimator for binary treatment is shown to be the effective influence function, it is unclear whether it still holds for continuous treatments without explicit computation. To verify, we calculate the influence function.
**Q2.** Why does the projected residual mean squared error make sense? Does it mean $\hat{q}$ is consistent for $q_0$?
**A.** The projected residual mean squared error measures how much $\hat{q}$ violate Eq. (3). This performance metric holds significant prominence in the theoretical analysis of minimax problems, as noted in references [1, 2, 3]. When the measures of the ill-posedness of inverse problems are bounded, $\hat{q}$ is consistent for $q_0$. Please refer to Remark 6.3 for details.
**Q3.** Please make the motivation clear about why this paper does not focus on the inference issue about the analysis of asymptotically distribution for a causal effect.
**A.** Currently, we cannot obtain the asymptotic normality due to the error introduced in kernel approximation. With this error, we can show in Theorem E.9 that our estimator is $n^{2/5}$-consistent, while the asymptotic normality means $\sqrt{n}$-consistent. Besides, according to [4], since the estimand is non-regular, therefore it may not enjoy the properties of $\sqrt{n}$-consistent and asymptotically normality. We illustrate this point through an empirical study in Appendix E.5.
**Q4.** The empirical coverage probability is not given. MSE only contains both bias and variance terms, which may not display the true statistic estimation accuracy and precision.
**A.** Thank you for your suggestions. We provide in Theorem E.9 that our estimator is $n^{2/5}$-consistent, which means with high probability, the error is $O(n^{-2/5})$.
**Q5.** This paper seems to consider the high-dimensional setting in experiments, what about theoretical properties?
**A.** In this context, we follow the setting in [5,6], in which the term "high-dimensional" merely indicates a scenario with relatively more covariates compared to those in section 7.1.1, without implying that the number of samples is smaller than the number of features. Therefore, our theory still applies to this case.
[1] Dikkala, Nishanth, et al. "Minimax estimation of conditional moment models." Advances in Neural Information Processing Systems 33 (2020): 12248-12262.
[2] Ghassami, AmirEmad, et al. "Minimax kernel machine learning for a class of doubly robust functionals with application to proximal causal inference." International Conference on Artificial Intelligence and Statistics. PMLR, 2022.
[3] Qi, Zhengling, Rui Miao, and Xiaoke Zhang. "Proximal learning for individualized treatment regimes under unmeasured confounding." Journal of the American Statistical Association (2023): 1-14.
[4] Colangelo, Kyle, and Ying-Ying Lee. "Double debiased machine learning nonparametric inference with continuous treatments." arXiv preprint arXiv:2004.03036 (2020).
[5] Xu, Liyuan, Heishiro Kanagawa, and Arthur Gretton. "Deep proxy causal learning and its application to confounded bandit policy evaluation." Advances in Neural Information Processing Systems 34 (2021): 26264-26275.
[6] Kompa, Benjamin, et al. "Deep learning methods for proximal inference via maximum moment restriction." Advances in Neural Information Processing Systems 35 (2022): 11189-11201.