Optimal Treatment Regimes for Proximal Causal Learning

A common concern when a policymaker draws causal inferences from and makes decisions based on observational data is that the measured covariates are insufficiently rich to account for all sources of confounding, i.e., the standard no confoundedness assumption fails to hold. The recently proposed proximal causal inference framework shows that proxy variables that abound in real-life scenarios can be leveraged to identify causal effects and therefore facilitate decision-making. Building upon this line of work, we propose a novel optimal individualized treatment regime based on so-called outcome and treatment confounding bridges. We then show that the value function of this new optimal treatment regime is superior to that of existing ones in the literature. Theoretical guarantees, including identification, superiority, excess value bound, and consistency of the estimated regime, are established. Furthermore, we demonstrate the proposed optimal regime via numerical experiments and a real data application.

Paper

References (77)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer of5M5/10 · confidence 5/52023-07-02

Summary

The authors present a new optimal individual treatment regime (ITR) within the proximal causal inference framework, which avoids the strong assumption of no unmeasured confounding. Instead, one assumes the effect of the unmeasured confounders flows exclusively through proxy variables, as defined through outcome-inducing and treatment-inducing confounding bridges. Compared to prior work, this optimal ITR that is defined with respect to a more flexible function class that depends on known confounders X, treatment-inducing confounding proxies Z, and outcome-inducing confounding proxies W.

Strengths

The proposed ITR is a natural extension of existing ITRs by using a function \pi(x) that selectively chooses between two existing ITRs based on known confounders x. Under the proximal causal inference frameowrk, the proposed ITR is proven to be superior to existing ITRs in the literature. (Existing ITRs from Qi 2023 can be viewed as special cases of the proposed ITR.) The authors introduce a simple plugin estimator for the proposed ITR and show that the value of the resulting estimator is determined by approximation error of \pi and the gain from using \pi. Simulation studies show that the proposed ITR is either superior or comparable to existing ITRs. The manuscript is clearly written. The authors provide a nice review of prior work in this area and clearly describe how their work builds on existing work.

Weaknesses

1. The proposed extension of the ITR function class appears quite incremental. The value of the proposed ITR follows directly from application of the tower rule. The paper would be greatly strengthened if the authors can show that this is the best one can do, e.g. showing that the value of a more complex ITR function class would be unidentifiable without much stronger assumptions. 2. In the simulation studies, the improvement in mean value when using the proposed optimal ITR over existing ITRs is large only in scenario 2. In all other scenarios, the improvement is small. Can the authors explain the behavior in this simulation study? Also, can the authors explain settings in which the proposed ITR is expected to substantially improve over existing ITRs? My guess is that the gain is biggest when (i) there are large differences between expected value at each X for the ITR with domain (X,W) and the ITR with domain (X,Z) and (ii) the optimal pi function has high variance (e.g. pi(X) is equal to 1 half of the time). Does this correspond to scenario 2? 3. The authors perform a real-data analysis in Section 5, which illustrates how the proposed ITR is different from existing ITRs. However, the authors do not calculate the values of the estimated ITR, so readers cannot compare the performance of the proposed ITR against existing ITRs. Do the authors have estimates of the values of the estimated ITRs? 4. The number of treatment-inducing confounding proxies and outcome-inducing confounding proxies were small in both the simulation studies and real-world data analysis. However, the practical appeal of the proximal causal inference framework is its use of proxies for unmeasured confounders, which would suggest the use of many variables as potential proxies. Can the authors include simulations that reflect more realistic settings where more proxy variables are used? How does the proposed method perform as the number of these proxies increases?

Questions

1. How does one do statistical inference for the value of the estimated ITR?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

The authors have not discussed limitations of the work.

Reviewer 193J7/10 · confidence 4/52023-07-05

Summary

The goal is to learn an optimal individual treatment rule (ITR) where the data suffer from unobserved confounding but where the researcher has a treatment proxy and an outcome proxy. While the general problem has been studied before by Qi et al (JASA 2023), this paper’s contribution is to broaden the class of ITRs. For a broader class of ITRs, the authors identify the value function and show that it exceeds the value function of the narrower class.

Strengths

My comments are brief because this is a strong paper. Originality: The essence of the improvement is that, for different covariate values x, one may either use the “outcome” ITR or the “treatment” ITR. This departs from previous work, where either the “outcome” ITR or the “treatment” ITR is used across covariate values. Quality: The proofs look correct, and the results are easy to interpret. Rates for the objects in Proposition 1 would be an improvement; see the question below. Clarity: The paper is well written and well referenced. Significance: The paper contributes to two popular literatures: proxies and ITRs. While its theoretical contribution is modest, it does appear to have practical relevance.

Weaknesses

The theoretical contribution is somewhat incremental. Take a pass to fix typos, e.g. “netwrok” on line 73. An extra sentence in Remark 1 would be welcome, that explains the point summarized as “originality” above.

Questions

Is it possible to derive rates for K(pi_hat) and G(pi_bar) in Proposition 1? Or can the authors at least pose this question for future work and give citations of where similar results are derived?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

Yes

Reviewer NTDa6/10 · confidence 2/52023-07-06

Summary

The paper discusses the optimization of treatment rules in the context of observational data and under assumptions of proximal inference. Various theorems are introduced, and a real data analysis performed using a healthcare example.

Strengths

Below is a list of perceived weaknesses. The paper is overall sound and the topic of importance. I appreciate the presence of the real data application. Assumptions and results clearly stated.

Weaknesses

Below is a list of perceived weaknesses. It was not clear to me how the empirical results compare to competing methodological baselines from other approaches (I don't believe the different values presented in the figure represent different algorithmic approaches). The paper is quite heavy on notation and, at least to me, light on intuitive explanation for findings as they are discussed, limiting insight into the inner workings of why the method works. I don't know the proximal causal inference literature well so am not well-positioned to discuss the contribution in that subfield of causal inference. I don't see a discussion of uncertainty estimation in the theoretical or empirical results. Uncertainty estimation in optimized treatment effect regimes can be difficult (e.g., the bootstrap may not be appropriate or may have poor coverage) but may be important to usefulness in practice.

Questions

Is there quantitative evidence that "proximal causes abound in real life scenarios"? The matter would seem dependent on substance area, or knowing this would seem to require access to nature's true DAG.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

I see no ethnical limitations here.

Reviewer VJQR6/10 · confidence 4/52023-07-07

Summary

Most estimation methods for individualized treatment rules (ITRs) assume no unmeasured confounders for valid causal inference. However, such an assumption can be unreasonable, such as when estimating ITRs from observational data. Previous work has applied proximal causal inference to estimate ITRs when this assumption is violated, but is restricted to policy classes that either exclude treatment-inducing confounding proxies or exclude outcome-inducing confounding proxies [1]. To this end, the authors propose estimating a stochastic mixture of both policy classes from [1] to yield a more flexible ITR. Theoretical and simulation results demonstrate the superiority of the proposed method.

Strengths

The assumption of no unmeasured confounders is nearly ubiquitous across ITR estimation methods, despite frequent violations when dealing with observational data. This makes the problem the authors are trying to solve - estimating ITRs when this assumption is violated - very significant. Moreover, the theoretical and simulation results are of sufficient quality to convince me that the method outperforms [1], the existing state-of-the-art in proximal learning.

Weaknesses

While I believe the merits and potential contribution of the paper outweigh its limitations, the theoretical and empirical results of the paper are weaker than that of previous work, and the clarity of the paper can be improved. I go into more detail below: 1. **Theoretical guarantees are much weaker than those of previous methods.** Convergence rates, finite-sample error bounds or asymptotic normality is often derived for ITR methods [2-5], including the method this work seeks to improve over [1]. However, while the authors prove that the proposed method will converge to a policy with better value than that of [1] asymptotically, they do not derive a rate of convergence or establish any finite-sample error bounds. Moreover, the asymptotic analysis from Appendix G assumes convergence of several estimators in $L\_\infty$ space. This is a much stronger assumption than the assumptions made in previous work, which only assumes convergence of estimators in $L\_2$ space (e.g. see Assumption 12 from [1] or Condition B5 from [3]). 2. **The real data analysis does not strengthen the validity of the proposed method.** When applying ITR estimators to real datasets, it is common practice to assess performance by using either (1) an estimator of the value [2-6] or (2) arguments based on domain knowledge that support the validity of the proposed method [1,6]. For example, in [1], the authors argue that the estimated policy is accurate by demonstrating that the estimated coefficients and interpretation of the policy is consistent with findings from previous literature. In contrast, the only conclusions that these authors draw from their real data analysis is that the proposed estimator gives different results than previously proposed methods. Such a conclusion says little about the validity or superiority of the proposed method. One way this analysis could be greatly improved is to look at the patients from Figure 3 where the recommended treatment differs between the proposed method and that of [1], and use domain knowledge or previous literature to argue that the recommendations given by the proposed method is more sound. Alternatively, the authors could make this conclusion by comparing the coefficients of the proposed policy with that of [1]. 3. **Empirical comparisons were relatively limited.** The authors only benchmark the proposed method against those of a single previous work, [1]. To conclude that the proposed method achieves state-of-the-art performance, the authors should benchmark the proposed method against additional baselines as well. For example, there are many methods that assume no unmeasured confounders which the authors could evaluate to demonstrate the utility of using a proximal causal inference framework (e.g. [7,8]). There are also other methods that relax the no unmeasured confounders assumption or have shown robustness when such assumptions are violated, such as instrument variable (IV) methods [3,9] and M-learning [4]. How are the assumptions made by proximal learning less restrictive than those made by IV approaches, and can such methods be applied to the simulated datasets? If so, the authors should benchmark the proposed method against these method. And if these methods are not applicable, the authors should explain why in the paper. In addition to the number of competing baselines being limited, the simulated datasets from this work all have the same sample size and behavior policy. When deriving new ITR methods, it is common practice to evaluate the method on datasets of different sample sizes (and if observational data is of interest, varying behavior policies) so as to demonstrate robust performance [2-5,7-11]. 4. **The implementation uses very simple estimators.** In theory, $d_z,d_w$ and $\delta$ could be estimated by any weighted classification and regression algorithm. However, in their empirical experiments, the authors only explore estimating $d_z$ and $d_w$ with linear policies and estimating $\delta$ with a Nadaraya-Watson estimator where the bandwidth is chosen based on a heuristic. Moreover, while $h$ and $g$ were estimated using neural networks, the architecture and number of training iterations was fixed a priori. This contrasts to previous works which use more cutting-edge machine learning approaches to estimate the ITR, such as kernel machines, random forests or neural networks, and adopts hyperparameters to the data at-hand using cross-validation [8]. Such works are especially prevalent in top-tier ML conferences [10,11], and better demonstrate broad applicability and flexibility of the proposed method. 5. **Many parts of the paper need to be better written to avoid confusion and address some open questions.** For example, for the real data analysis, it is not clear what assumptions proximal learning is making and how it is useful for the analyzed dataset. It is mentioned that patients were arranged in a "control group" on line 289. Were patients randomized to receive a treatment? If so, wouldn't the no unmeasured confounders assumption hold, as treatment assignment is not being affected by any unmeasured covariates? Also, what is the logic behind the choice of $Z$ and $W$ on line 296 (e.g. why are the variables in $Z$ expected to affect treatment but not outcome), and what specifically are the unmeasured confounders $U$ that we are trying to adjust for? Finally, it is stated that the outcome is censored. Does this mean that 30-day survival rate is censored for some of the patients? It is well-known that optimizing censored outcomes without adjusting for the censoring mechanism can yield bias [5]. >> Here are some other suggestions to reduce points of confusion and improve readability: a. The explanation of how the proposed method improves upon [1] in the Introduction section (lines 56-65) is confusing. For example, it is stated on line 69-61 that [7] maps an ITR with domain $\mathcal X\times\mathcal W\times\mathcal Z$ with the domain being restricted to $\mathcal X\times\mathcal W$ to $\mathcal X\times\mathcal Z$. While this makes more sense after reading section 2.2, these statements initially appear contradictory. Also on lines 59 and 63 "two optimal in-class ITRs" should be changed to "these two optimal in-class ITRs" to clarify that the authors are referring to the classes mentioned on line 57. b. The paper has many typos. For example, "Tchetgen Tchetgen et al" in line 34 should include the year and a link to the reference, "netwrok" on line 73 should read "network", and on line 169 the authors should add $V(d_{z\cup w})$ to the argmax. c. On line 142-143 the authors state that Assumption 3 assumes "Z has sufficient variability with respect to the variability of U". But isn't assumption 3 actually assuming that U has sufficient variability with respect to the variability of Z? d. On line 173 it is stated that $\mathbb E[Y(a)|X,U]$ may not be identifiable under proximal causal inference. But doesn't assumptions 1-5 allow for such identifiability? e. Remark 1 is not true. For example if $\pi(X)=0.5$ then $\pi$ is constant but the policy class will not be in $\mathcal D_{\mathcal Z}\cup \mathcal D_{\mathcal W}$. We actually need the restriction that $\pi$ is both constant and in the set $\\{0,1\\}$. f. The results of Appendices E and G should appear in the main paper as propositions, theorems or corollaries. g. For sections where over 10 references are cited back-to-back, I think readability would be improved if these citations appeared in chronological order. h. The authors should add a Discussion section that summarizes the results of the paper and proposes important avenues for future work (also see my comments on the Limitations section). References: 1. Qi Z, Miao R and Zhang X. Proximal Learning for Individualized Treatment Regimes Under Unmeasured Confounding. JASA. 2023. 2. Zhao Y, Zeng D, Rush AJ and Kosorok MR. Estimating Individualized Treatment Rules Using Outcome Weighted Learning. JASA 107 (499): 1106-1118. 2012. 3. Qiu H et al. Optimal Individualized Decision Rules Using Instrumental Variable Methods. JASA 116 (533): 174-191. 2021. 4. Wu P, Zeng D and Wang Y. Matched Learning for Optimizing Individualized Treatment Strategies Using Electronic Health Records. JASA 115 (529): 380-392. 2020. 5. Zhao YQ, Zeng D, Laber EB, Song R, Yuan M and Kosorok MR. Doubly robust learning for estimating individualized treatment with censored data. Biometrika 102 (1): 151-168. 2015. 6. Raghu et. al. Continuous State-Space Models for Optimal Sepsis Treatment: a Deep Reinforcement Learning Approach. MLHC 2017. 7. Zhao YQ, Laber EB, Ning Y, Saha S and Sands BE. Efficient Augmentation and Relaxation Learning for Individualized Treatment Rules using Observational Data. JMLR 20: 1-23. 2019. 8. Zhou X, Mayer-Hamblett N, Khan U and Kosorok MR. Residual Weighted Learning for Estimating Individualized Treatment Rules. JASA 112 (517): 169-187. 2017. 9. Pu H and Zhang B. Estimating optimal treatment rules with an instrumental variable: A partial identification learning approach. JRSS-B 83 (2): 318-345. 2021. 10. Yoon J, Jordon J and van der Shaar M. GANITE: Estimation of Individualized Treatment Effects using Generative Adversarial Nets. ICLR 2018. 11. Chen Y, Zeng D, Xu T and Wang Y. Representation Learning for Integrating Multi-domain Outcomes to Optimize Individualized Treatment. NeurIPS 2020.

Questions

Questions that I feel the paper should address and general weaknesses of the paper can be found in the Weaknesses section. As to my suggestions, it may not be feasible to address weaknesses (1) and (4) prior to the camera-ready. However, I would suggest the authors at least acknowledge weakness (1) as a limitation, and discuss how to implement their method in a more data-adaptive manner, even if they don't change the implementation for the empirical experiments. Such a discussion should include how to tune hyperparameters for all functional classes, including estimators of $d_z$, $d_w$, $\delta$, $h$ and $g$. As to weaknesses (2), (3) and (5), I think it would be feasible to mostly address these weaknesses prior to the camera-ready, and my suggestions on how to address them can be found in the Weaknesses section.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

I did not notice the authors acknowledge any limitations of the proposed work. I would recommend adding a "Discussion" section that summarizes the results of the work, and addresses limitations by discussing important avenues for future work. In addition to what was discussed in the Weaknesses and Questions sections, another important limitation is that the proposed method assumes that assumptions (1)-(5) hold and that both $h$ and $g$ are correctly specified, while $d_z$ and $d_w$ appear to only require some of these assumptions to hold.

Reviewer VJQR2023-08-15

Response to "Rebuttal by Authors"

I appreciate the authors' thorough rebuttal. Some of my concerns about the paper have been addressed, though other concerns still remain. I go into more detail below: 1. I still feel that the lack of convergence rates and finite-sample error bounds, and the requirement of $L_\infty$ convergence, make the theoretical results of this work weaker than that of previous work. While the authors proved that their algorithm is consistent in their rebuttal, it seems that their proof also assumes $L_\infty$ convergence via reliance on assumption 6 and Appendix G, and I do not see how this directly leads to a finite-sample error bound for the estimated regime. Moroever, they state that the $L_\infty$ convergence assumption is required to "guarantee that for each $X$, the conditional expected values based on $\hat h$ and $\hat g$ are consistent, which is not required by previous literature." However, [1] also involved estimating similar quantities and managed to avoid $L_\infty$ convergence assumptions. Therefore, I feel it would be possible to avoid such an assumption here as well. 2. While the table of empirical value estimates provided in Table 1 of the authors' rebuttal partially alleviates my concerns regarding the real data analysis, I am still concerned with the way in which performance is being assessed. Specifically, how is value being estimated in Table 1? If the authors are using importance-sampling estimators, these estimators are biased in the presense of unmeasured confounders, which is exactly what this paper is trying to address. Morevoer, if they are using different estimators for different methods, than this makes it difficult to compare value estimates between methods. Therefore, I feel the real data results would be further strengthened by performing a qualitative analysis of the regime estimated by the authors' proposed method and using domain knowledge to validate its performance, similar to what has been done in previous work [1,2,3]. For example, one way to do this is to compare estimated coefficients or treatment recommendations between the proposed method and that of [1], and argue that the proposed method's estimates are more consistent with previous literature. 3. Between the added comparisons involving additional baselines in Figure 1 of their rebuttal, their clarification on why other baselines such as those based on instrument variables would not be applicable, and their promise to add additional experiments with varying sample sizes and behavioral policies in the revision, all of my concerns related to the weaknesses of the simulated experiments have been satisfied. One remaining suggestion is to clarify in the revision why methods based on IV variables are not applicable to the simulated datasets. 4. I appreciate the authors promising to add discussion on how to combine their methodology with ML algorithms, though I would have liked to see the authors provide such a discussion in their rebuttal. For example, while it is straightforward to apply ML to estimation of $d_z,d_w$ and $\delta$, including tuning hyperparameters, I am curious how the authors would propose tuning hyperparameters to the data at hand when estimating $h$ and $g$, including selecting the network architecture. 5. I appreciate the authors' clarifications regarding the real data set and provided references. I should note that reading [1] did not alleviate my confusion regarding the real dataset. Therefore, I recommend that the authors expand upon the provided clarifications in the revision so that future readers do not suffer from the same confusion. I also appreciate the authors clarifying other points of confusion and promising to add discussion of limitations in the revision. References: 1. Qi Z, Miao R and Zhang X. Proximal Learning for Individualized Treatment Regimes Under Unmeasured Confounding. JASA, 2023. 2. Raghu et. al. Continuous State-Space Models for Optimal Sepsis Treatment: a Deep Reinforcement Learning Approach. MLHC 2017. 3. Luckett DJ et. al. Estimating Dynamic Treatment Regimes in Mobile Health Using V-Learning. JASA, 2019.

Authorsrebuttal2023-08-16

We sincerely express our gratitude to the reviewer for their thorough examination of our rebuttal and for generously sharing invaluable insights. We are genuinely appreciative of the reviewer's acknowledgment of our endeavors in addressing the comments in W3 and W5. We are devoted to enhancing our text in the revision as suggested. To address the reviewer's raised concerns regarding W1, W2, and W4, we will now provide responses below. Response to 1: In contrast to the approach taken by Qi et al. (2023, JASA), whose Theorem 5.1 highlights the consistency of $V(\hat{d}z)$ and $V(\hat{d}w)$, our focus, as articulated in our rebuttal, is distinct: to make the value of our proposed regime consistent, we strive to ensure the consistency in the expected values for each stratum $X$, i.e., the consistency of a conditional value function, which is not required by Qi et al. (2023, JASA) or other works within this framework. Therefore, we impose the $L_{\infty}$ convergence assumption. We emphasize that the consistency result we have developed is quite novel as the distinct phenomenon of our problem is brand new and different from the previous literature. Regarding the derivation of convergence rates and finite sample error bounds, we recognize that achieving such results is a non-trivial endeavor. Bounding the error metric and establishing advanced theory has captured our keen interest. We will include the discussion of limitations and future work in the revision. Response to 2: We thank you for your insightful suggestion. Our chosen performance evaluation metrics take the latter approach, i.e., using different estimators for different methods. Admittedly, employing such criteria may potentially dilute the persuasiveness of comparisons between methodological value estimates. However, this evaluation is fair to each baseline method, and gives an unbiased estimation under each identification strategy. Regarding the inclusion of qualitative analysis, we present an illustrative example below, with plans to expand this discussion in the revision. Notably, the coefficient of cat1_lung is negative with a minor magnitude for $\hat{d}_w$, contrasting with a positive and relatively large coefficient observed for $\hat{d}_z$. These outcomes mirror those outlined in Qi et al. (2023, JASA). This finding suggests that, within the primary disease category of patients with lung cancer, $\hat{d}_z$ advocates for undergoing RHC, while $\hat{d}_w$ displays a notably inconclusive trend. As evidenced by $\hat{\pi}$, the prevailing trajectory for patients with cat1_lung = 1 involves a strong inclination toward undergoing RHC, i.e., $\hat{\pi}=1$, aligning with the guidance offered by $\hat{d}_z$. Significantly, the domain knowledge underscores the potential for patients with advanced lung cancer to develop complications like pulmonary hypertension and coma, potentially warranting RHC for assessing pulmonary vascular changes and informing treatment strategies (Galie et al., 2009). This body of domain-specific knowledge lends support to the recommendations offered by our proposed regime. Response to 4: We appreciate your attention to hyperparameter tuning and architectural selection during the estimation of bridge functions. Within our experiment, we have diligently conducted both architecture search and hyperparameter optimization, exploring a comprehensive range of configurations. Generally, we consider estimating bridge functions employing multilayer perceptrons with 2-8 fully connected layers with a variable number of hidden units. We then perform a grid search over the following parameters: learning rate (3e-3, 3e-4, 3e-5), $L_2$ penalty coefficient (1e-3, 1e-2, 1e-1), number of epochs (100, 150, 200, 250), batch size (100, 250, 500), depth of network (2, 4, 8), width of network (10, 40, 80). For every permutation of these parameters, we train a network based on the determined architecture and parameter values. Subsequently, we compute the empirical risk, as detailed in Appendix J. Our aim is to pinpoint the parameter combination that yields the lowest empirical risk. These identified optimal parameters are then utilized to construct a refined neural network, which, in turn, serves as the foundation for conducting estimations. For a comprehensive understanding of these intricacies, we suggest delving into supplementary Section B in Kompa et al. (2022, NeurIPS). This supplementary resource offers detailed insights into the specific hyperparameter choices and architectural dimensions. References: Galie N, Hoeper M M, Humbert M, et al. Guidelines for the diagnosis and treatment of pulmonary hypertension: the Task Force for the Diagnosis and Treatment of Pulmonary Hypertension of the European Society of Cardiology (ESC) and the European Respiratory Society (ERS), endorsed by the International Society of Heart and Lung Transplantation (ISHLT)[J]. European heart journal, 2009, 30(20): 2493-2537.

Reviewer VJQR2023-08-17

Response to "Official Comment by Authors"

I thank the authors for elaborating on hyperparameter tuning strategies, and more generally on trying so hard in good faith to address these concerns. Between their initial rebuttal and their most recent reply, many (thought not all) of my concerns now been addressed, and thus I will raise my score of the paper from "borderline accept" to "weak accept". I still have a few remaining suggestions to the authors, which I detail below: Regarding W1: I believe I now better understand why the authors are making an $L_\infty$ assumption: They are using it to prove consistency of the conditional value function, which is stronger than proving marginal value consistency from Qi 2023. Nonetheless, it stands that without an $L_\infty$ assumption, Qi 2023 still at least achieves marginal value consistency, while this work does not. For the revision, I recommend the authors better explain why they are making an $L_\infty$ assumption in contrast to previous work. I also recommend they try to achieve similar consistency results to Qi 2023 with an $L_2$ assumption, provided doing so would not be that difficult. Regarding W3: The qualitative analysis provided by the authors is more in-line with what I was hoping for. However, rather than compare the proposed method to $\hat d_w$ using domain knowledge, I recommend the authors instead compare the proposed method to $\hat d_{z\cup w}$ for the revision. Also, when discussing the value table for the real data analysis, I recommend the authors explaining that each method is being evaluated with different methods, and acknowledging that this is a limitation of the comparison.

Authorsrebuttal2023-08-18

We express our sincere gratitude to the reviewer for elevating the score of our work and giving invaluable suggestions that have provided us with a clear direction for our revision. For W1, we will add explanations and limitations of utilizing the $L_\infty$ assumption. For W3, we were comparing the proposed method to $\hat d_z$ and $\hat d_w$ using domain knowledge. For this dataset, note that $\hat d_{z \cup w}$ and $\hat{d}_{z}$ lead to the same regime. Hope this clarifies the concern. Besides, we will acknowledge that each method is being evaluated with different methods, and the values might not be fully comparable when some of the assumptions are not satisfied.

Reviewer VJQR2023-08-20

Response to "Official Comment by Authors"

Following up on my previous point on W3, I meant it would make the revision stronger to demonstrate that your proposed method is more consistent with domain knowledge and previous literature than $\hat d_{z\cup w}$, instead of only showing it is superior to $\hat d_w$ as is done in your previous comment.

Authorsrebuttal2023-08-21

Thank you. In the previously presented example, we demonstrated the advantage of $\hat d_z$ for a subgroup of patients with cat1_lung = 1. However, it's important to note that the whole group of patients can be regarded as unions of multiple subgroups based on various distinct features, and the superiority of $\hat d_w$ is evident in some subgroups (e.g., amihx). This shows that our proposed approach offers superior efficacy compared to $\hat d_{z \cup w}$ as our methodology incorporates selection through $\pi$. We will include the discussion in the revision and hope this explanation will address your concern.

Reviewer of5M2023-08-18

Thank you for the response

Thank you for the response. I think the work is now stronger with these additional derivations and experiments. My main concern is that the model class is still somewhat limited, and I will keep my score as is.

Authorsrebuttal2023-08-18

We thank the reviewer’s thorough evaluation and recognition of our response, as well as for sharing insightful suggestions.

Reviewer NTDa2023-08-21

Response to Author Comment

Thank you to the authors for their comments. The various modifications seem to have improved the piece, so I raise my score by one. To bolster the impact of the work, I would emphasize again the importance of comparative baselines against existing methods (to address "even if method X, Y, or Z use different assumptions, perhaps they do better" type questions), uncertainty estimation, and providing clear guidance for applied researchers about how to "know" if they're in the proximal regime.

Authorsrebuttal2023-08-21

We appreciate the reviewer's recognition of our response and valuable suggestions. We will enhance baseline method analysis and discuss uncertainty estimation as well as application guidance in the revision.

Reviewer 193J2023-08-21

I appreciate the response. I continue to recommend the high rating.

Authorsrebuttal2023-08-21

We thank the reviewer’s acknowledgement of our work and the response.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC