Temporal Causal Mediation through a Point Process: Direct and Indirect Effects of Healthcare Interventions

Deciding on an appropriate intervention requires a causal model of a treatment, the outcome, and potential mediators. Causal mediation analysis lets us distinguish between direct and indirect effects of the intervention, but has mostly been studied in a static setting. In healthcare, data come in the form of complex, irregularly sampled time-series, with dynamic interdependencies between a treatment, outcomes, and mediators across time. Existing approaches to dynamic causal mediation analysis are limited to regular measurement intervals, simple parametric models, and disregard long-range mediator--outcome interactions. To address these limitations, we propose a non-parametric mediator--outcome model where the mediator is assumed to be a temporal point process that interacts with the outcome process. With this model, we estimate the direct and indirect effects of an external intervention on the outcome, showing how each of these affects the whole future trajectory. We demonstrate on semi-synthetic data that our method can accurately estimate direct and indirect effects. On real-world healthcare data, our model infers clinically meaningful direct and indirect effect trajectories for blood glucose after a surgery.

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer 5ke67/10 · confidence 4/52023-07-06

Summary

This paper studies how to estimate the direct and indirect effctes of healthcare interventions. The general idea of this paper is to model the mediation process and outcome process jointly. More specifically, it considers the mediation process as a temporal point process conditioned on the past mediation, outcome and treatment data. It allows two causal paths: direct path models the direct effect of treatment and indirect path models the path treatment->mediation->outcome. The authors prove that under their three assumptions, the two effects can be represented by two terms in their non-paraametric temporal point process model. Experimental results shows the advantage of their proposed model.

Strengths

(1) The proposed model studies an interesting problem: how to distinguish indirect and direct effect, which is important in healthcare (2) The overall design of their causal model are reasonable. Althought their assumptions are not easy to verify on the data, they have tried their best to give convincing analysis to the data. (3) The motivation is clear and convincing.

Weaknesses

(1) Did you consider that in real world, the treatment may be correlated with the outcome, leading to bias in the model? Did you try to reduce the issue with IPTW or other method to debias? (2) Some recent related works about handling treatment effect and causal inference with non-parametric temporal point process model are missing. For example, [1] studies the treatment effect in healcare, too. [2] considers how to debias the neural temporal point process in the context of social media analysis. [3] studies how to sample the counterfactual sequences from temporal point process. [1] Gao, Tian, et al. Causal Inference for Event Pairs in Multivariate Point Processes. NeurIPS 2021 [2] Zhang, Yizhou, et al. Counterfactual Neural Temporal Point Process for Estimating Causal Influence of Misinformation on Social Media. NeurIPS 2022. [3] Noorbakhsh, Kimia , and M. G. Rodriguez . Counterfactual Temporal Point Processes. NeurIPS 2022.

Questions

Did you consider that in real world, the treatment may be correlated with the outcome, leading to bias in the model? Did you try to reduce the issue with IPTW or other method to debias?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

They have discussed.

Reviewer NnCo5/10 · confidence 4/52023-07-07

Summary

The paper aims to estimate the direct and indirect treatment effects of healthcare interventions. The authors model the mediator as a point process and propose a non-parametric mediator–outcome model where the mediator is assumed to be a temporal point process that interacts with the outcome process. The authors conduct experiments on a real-world RCT dataset and a semi-synthetic dataset. The experiments on the synthetic dataset show that the proposed model outperforms the baselines on the treatment effect estimation tasks.

Strengths

- The authors propose a new approach to model the direct and indirect effects of healthcare interventions. - The authors conduct experiments on both real-world and semi-synthetic datasets. The results demonstrate that the proposed model outperforms the baselines. - The implementation code is available.

Weaknesses

- Treatment A, as the confounder, affects both mediator M and outcome Y. M affects Y. The confounding bias would make the estimation E[Y|A,M] inaccurate. Figure 5 shows the difference of diet density pre- and post-surgery, which demonstrates the existence of confounding bias. Without consideration of the confounders, the direct and indirect treatment effect estimation could be inaccurate. - The authors claim that (A1, A2, A3) might not hold in observational studies and they are not statistically testable. I have the concern that if the assumptions do not hold, can the proposed model be applied to real-world applications? How to evaluate the potential risk of the model. - A1,A2,A3 are very strong assumptions. In real-world settings, the no-unobserved confounder assumption may not hold, so it is necessary to conduct a sensitivity analysis of how sensitive or robust the proposed models are to the unobserved confounders. - Bariatric surgery could cause weight loss. Another mediator weight would significantly change after the surgery. - Some details are missing. The authors just use diet and surgery to predict blood glucose. Is any detailed information about the diet, like nutrients including sugar, and starch? Are patients’ demographics (weights, age, height) used in the experiments? - It would be better if the authors display the factual prediction performance, like MSE for glucose prediction. - It is unclear what the variables m and o in Eq. (9) mean.

Questions

See above.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

NA

Reviewer jr2N5/10 · confidence 3/52023-07-08

Summary

The paper defines direct and indirect effects in complex healthcare time-series as dynamic stochastic processes and theoretically provides causal assumptions for identifiability. This model allows for an external intervention influencing both mediator and outcome sequences simultaneously and captures time-delayed interactions among them.

Strengths

1. The authors proficiently present the estimated direct and indirect effects as longitudinal counterfactual trajectories, along with the requisite theoretical causal assumptions for their identification. 2. The method proposed is neat and articulated with good clarity.

Weaknesses

1. The method's scalability to larger datasets poses a concern. There exists many methods specifically designed for high-dimensional mediation analysis in time series [1][2][3]. It would be enlightening to observe how the proposed method compares to these in handling complex datasets. 2. To my understanding, Figure 1a may not accurately depict Zeng et al. 2021 [4]. It seems the original work allows for past mediators to have an influence on future outcomes. References [1] Chén, Oliver Y., et al. "High-dimensional multivariate mediation with application to neuroimaging data." Biostatistics 19.2 (2018): 121-136. [2] Zhang, Haixiang, et al. "Mediation analysis for survival data with high-dimensional mediators." Bioinformatics 37.21 (2021): 3815-3821. [3] Luo, Chengwen, et al. "High-dimensional mediation analysis in survival models." PLoS computational biology 16.4 (2020): e1007768. [4] Zeng, Shuxi, et al. "Causal mediation analysis for sparse and irregular longitudinal data." The Annals of Applied Statistics15.2 (2021): 747-767.

Questions

1. What advantage does the utilization of a marked point process offer in modeling mediators compared to the approach in [1], where the observed mediator is considered as drawn from a smooth underlying process? 2. Does the proposed method allow for modeling a high-dimensional observed mediator? 3. In line 119-120, the mediator process M considers the number of occurrences of the mediating event up until time $\tau$ and the value of the mediator at time $\tau$. I'm curious if the occurrence count alone is sufficient to model the process, or if the past values of the mediator should also be considered? 4. Regarding line 146, how is the continuation of the outcome after the intervention at $t_a$ represented? Should it be an equal average over all timepoints post $t_a$, or should it be a weighted average giving more importance to timepoints closer to $t_a$? 5. For quick clarification, in line 159, $H_{\leq \tau}$ refers to the history up until time $\tau$. Does this history include both mediator and outcome? Reference [1] Zeng, Shuxi, et al. "Causal mediation analysis for sparse and irregular longitudinal data." The Annals of Applied Statistics15.2 (2021): 747-767.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Yes, the authors addressed their limitation, which involve untestable causal assumptions and scalability of the method to larger data sets.

Reviewer bzqS7/10 · confidence 3/52023-07-26

Summary

The paper's outstanding qualities lie in its well-articulated presentation and its precise experimental design. It amalgamates the earlier research findings of [Zeng et al., 2021] and [Hızlı et al., 2022] with the innovative notions put forth by [Robins et al., 2022] on indirect effects. The authors tackle pragmatic issues, such as the effects of surgery on a patient's blood sugar levels in relation to their diet, by formulating pertinent questions. For instance, they question whether optimal post-surgery mediation can entirely regulate blood sugar levels, or if there exist uncontrollable surgery-induced effects on blood sugar levels that resist management through mediation and diet adjustments. In their methodological approach, the authors utilize non-parametric models of the temporal data-generating process. They employ a Marked Point Process (MPP) akin to [Hızlı et al., 2022] for meal intake (the mediating factor) and a non-parametric Gaussian Process (GP) for the outcome. The former is modeled as a combination of a counting process (number of meals) and a dosage process (carb intake per meal) using a non-parametric Poisson Process and the latter as a Gaussian process. The mediator model is trained to predict the mediator based on the intervention and the outcome model is trained with both the mediator and intervention given as input. This approach is put to the test in a series of synthetic experiments, validating its predictive strength against baseline models like Zeng et al., 2021. They further apply it in a real-world setting, successfully reproducing biological insights and potentially addressing the question of the degree to which surgery impacts diet changes. While at first glance, this paper might seem to bear similarities to [Hızlı et al., 2022], it stands out through its adept combination of the best methods for examining direct and indirect cause-and-effect relationships in a temporal setting. The paper, with its organized code and clear writing style, has the potential to become a valuable asset. However, I'm leaning towards accepting this paper on the condition that certain concerns regarding its originality and novelty are addressed, which I will detail in the following.

Strengths

1. The paper is commendable for its realistic problem setup, which is articulated in Section 5.1 dealing with corner cases where the assumptions might falter: the existence of hidden confoundings and the violation of assumptions A.1 to A.3. Such validation is crucial in causal studies to confirm assumptions and identify any unnoticed confoundings that may influence our conclusions. Furthermore, the experiments and predictions are consistent with the clinically significant direct and indirect impacts of bariatric surgery on blood glucose levels. 2. The approach to modeling the temporal dynamic is robust, anchoring its foundation on recent, proven work that adds to its credibility. 3. The paper's eloquent presentation is worthy of note. The reading experience is enhanced by effective use of color-coding to differentiate between mediator and direct interventions. A minor suggestion would be to consider adaptations for grayscale printed versions of the paper. For instance, the caption of Figure (2) includes light and dark blue color coding, which could be made more distinguishable by slightly altering the arrow patterns.

Weaknesses

1. The theoretical advancement of the study appears relatively marginal. While the exploration of direct vs. indirect causal effects in a temporal setting is engaging and the experiments provide valuable insights, I have some reservations about two of the claimed main contributions: * Dynamic causal mediation with a point process mediator: [Hızlı et al., 2022] have previously introduced point process mediator modeling. The novelty here is questionable, given that in the prior work, the treatment was the mediator itself. * A mediator-outcome model with an external intervention: The distinctiveness here is the training of two models: pre-intervention and post-intervention. However, the applied intervention is overly simplistic, offering limited theoretical innovation or insight. I have proposed, in the "Questions" section, the inclusion of the theory behind more complex interventions and experimentation on the simpler case. Yet, as it stands, this contribution mainly replicates the approach from [Hızlı et al., 2022], but uses two models to account for the intervention. 2. Table 1 presents results suggesting that the direct causal impact of surgery outweighs its indirect effects. Although the insights from 5.1.3 and 5.1.2 align with existing studies, there seems to be no supportive evidence for this hypothesis. Perhaps incorporating relevant literature explanations into the discussion would be beneficial. While the coherence between findings in 5.1.2 and 5.1.3 lend some validation to the model, it would still be advantageous to have literature support for 5.1.4.

Questions

1. Even though the paper presents a succinct and coherent narrative, it's hard to ignore that the methodology could easily extend to cases where the intervention itself is also a point process. For instance, one might consider a patient's long-term history and periodic clinical treatments. In such scenarios, $NIE$ and $NDE$ could be defined at different time points. It might be beneficial to include the theory behind this in the appendix section. The theoretical framework in sections 2 and 3 would work if one defines $N_A: [0, T] \to \mathbb{N}$, and instead of developing two distinct models, a more comprehensive model could be formulated that includes the history of interventions. This approach could also minimize the chance of future incremental papers being published. 2. What is the model's predictive power under model misspecification? Currently, the paper only presents results assessing predictive power in semi-synthetic scenarios where the model is appropriately specified. However, it would be beneficial to conduct experiments in scenarios reflective of real-world settings where model misspecification is common. Although the paper does incorporate a real-world setting, it is used merely for extracting qualitative observations. A comparative analysis of prediction power, similar to the semi-synthetic scenario but including model misspecification, is missing.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

2 fair

Limitations

1. Like any causality-oriented study, this paper relies on certain assumptions for its model to function effectively. However, the authors have acknowledged this constraint adequately. They have also indicated possible real-world scenarios with hidden confounders where these assumptions might be compromised. This demonstrates a good understanding of the model's limitations. 2. A further limitation of the method, as previously hinted, is its inability to handle more complex interventions. Currently, the model considers total effect, direct, and indirect effect only in relation to a unitary, binary intervention point process. It could potentially be enhanced by extending its capacity to accommodate non-binary scenarios (for different intervention styles, such as dosage treatments) or non-unitary processes (considering the entire electronic health record of a patient over a long period), which could provide a more comprehensive understanding.

Reviewer bzqS2023-08-16

I would like to thank the authors and acknowledge reading their response.

Reviewer bzqS2023-08-18

Considering the underlying theory of the extension, along with the added support for version 5.1.4, I find it justifiable to marginally increase my rating to a 7. While I still perceive the theory and methodology as somewhat incremental, the compelling practical results that back them up, paired with the comprehensive responses provided, lead me to believe that an increase in score is warranted.

Area Chair x3ao2023-08-18

Response to rebuttal

Dear reviewer, The author rebuttal appears to have presented several targeted responses to your questions. Are your questions appropriately addressed? If they are, would you consider re-assessing your score in light of them. If not, please do provide additional context and feedback to the author. In either case, please provide an acknowledgement of the effort the authors put in, why your questions have (or have not) been addressed and what your assessment of the work is in light of this evidence with a view to reach consensus with the other reviewers on this work. -AC

Area Chair x3ao2023-08-18

Response to rebuttal

Dear reviewer, The author rebuttal appears to have presented several targeted responses to your questions. Are your questions appropriately addressed? If they are, would you consider re-assessing your score in light of them. If not, please do provide additional context and feedback to the author. In either case, please provide an acknowledgement of the effort the authors put in, why your questions have (or have not) been addressed and what your assessment of the work is in light of this evidence with a view to reach consensus with the other reviewers on this work. -AC

Reviewer NnCo2023-08-18

Thanks for the comprehensive replies from the authors. I still have some concerns: Regarding the sensitivity analysis, it would greatly enhance the robustness of your study if you could also assess the model's performance when confounding influences impact both the training and test datasets. Given the realities of real-world scenarios, the presence of hidden confounders can significantly impact the test data, thereby warranting an exploration of this aspect. In relation to the factual prediction experiments, while I appreciate the utilization of a 2-day training data setup with the remaining day for validation, I would suggest a more patient-centric division of the training and test sets. By avoiding potential overfitting concerns due to shared patient data in both sets, your results would be further strengthened. Additionally, there is a need for clarification on whether cross-validation was conducted in the experiments presented in the original manuscript. If the experiment setup aligns with the factual prediction experiments outlined in the attached pdf, the persuasiveness of the outcomes may be compromised. It's worth highlighting that metabolic rates exhibit considerable variation among individuals, a phenomenon intricately tied to demographic factors such as age, gender, body weight, and height. These demographics could be potential confounding variables. Their absence could potentially introduce inaccuracies in the interpretation of treatment effects. Moreover, such demographic insights are integral to randomized controlled trials (RCTs), offering a lens to assess the representativeness of RCT data.

Authorsrebuttal2023-08-21

Thanks for the insightful comments. **More comprehensive sensitivity analysis study:** For the final version, we are happy to extend our experiment with confounding in both training and test sets. **A more patient-centric division of the training and test sets for the factual prediction experiments:** Our model is hierarchical: the baseline and response magnitude are individual-specific, and the response shape is shared. We have not tried predicting patients with models trained purely on other patients, but the accuracy would likely be low due to relatively large differences between individuals. Predictions tuned for an individual are well-motivated also from the application point-of-view, and the added experiment reflects this. **Clarification on cross-validation in the experiments:** The most important hyperparameters were the GP lengthscales, which were selected by a combination of domain knowledge and validation of the model fit. For meal-response, we considered lengthscales 0.15 h, 0.3 h, 0.5 h, 1.0 h, and for the baseline, lengthscales 5 h, 10 h, 20 h. We’ll clarify the treatment of these and other hyperparameters in the Appendix. **If the experiment setup aligns with the factual prediction experiments outlined in the attached pdf, the persuasiveness of the outcomes may be compromised:** After contemplating this since the first rebuttal, we felt like we failed to emphasize that the main contribution is really on estimating direct and indirect effects, suitable for answering counterfactual questions: for example, what would happen if only the diet changed but not metabolic processes, or vice versa. For such questions, the factual outcomes are not available and therefore we feel that an experiment comparing the MSE between predictions and factual outcomes is inevitably a bit misaligned with the rest of the paper. On the other hand, we feel a more justified validation is obtained by measuring the accuracy of direct and indirect effect estimates in the semi-synthetic experiment, and benchmarking the estimates with domain knowledge in the real-world study. We apologize for not communicating this clearly in the earlier response but hope this clarifies now. Nevertheless, we will also include the factual prediction experiment in the Appendix as promised. **Potential confounders:** We agree and will further emphasize this as a limitation of the present study.

Reviewer jr2N2023-08-18

Thank you for the comprehensive clarification on various points that I raised in my review. I appreciate the time and effort you put into addressing my concerns, and I find it justifiable to marginally increase my rating to a 5 and my confidence score to 3. I have one follow-up question on the choice of models. Since the method and problem formulation are agnostic to the choice of models, have you considered or conducted any empirical studies to validate the model's robustness across different model choices? Overall, I look forward to seeing the final version of this manuscript, as temporal causal mediation is a very intereting field with great application potential. However, I am really interested in how this might scale, especially since high-dimensional observed mediators are getting more and more prevalent in many applications. Exploring this could uncover even more impactful insights and opportunities.

Authorsrebuttal2023-08-21

Thanks for your thoughtful comment. **Different modeling choices:** With the synthetic datasets, we considered multiple choices for the mediator and outcome models (Table 2). In the real world case-study, we used our complete model, and in the new real-world results added in the rebuttal we also included alternatives, though neural ODE based approaches were not included in this study.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC