Identification and Estimation of the Bi-Directional MR with Some Invalid Instruments

We consider the challenging problem of estimating causal effects from purely observational data in the bi-directional Mendelian randomization (MR), where some invalid instruments, as well as unmeasured confounding, usually exist. To address this problem, most existing methods attempt to find proper valid instrumental variables (IVs) for the target causal effect by expert knowledge or by assuming that the causal model is a one-directional MR model. As such, in this paper, we first theoretically investigate the identification of the bi-directional MR from observational data. In particular, we provide necessary and sufficient conditions under which valid IV sets are correctly identified such that the bi-directional MR model is identifiable, including the causal directions of a pair of phenotypes (i.e., the treatment and outcome). Moreover, based on the identification theory, we develop a cluster fusion-like method to discover valid IV sets and estimate the causal effects of interest. We theoretically demonstrate the correctness of the proposed algorithm. Experimental results show the effectiveness of our method for estimating causal effects in bi-directional MR.

Paper

References (48)

Scroll for more · 36 remaining

Similar papers

Peer review

Reviewer UQEn6/10 · confidence 3/52024-07-12

Summary

The paper addresses the challenge of estimating causal effects in bi-directional Mendelian randomization (MR) studies using observational data, where invalid instruments and unmeasured confounding are common. It investigates theoretical conditions for identifying valid instrumental variable (IV) sets and proposes a cluster fusion-like algorithm to discover these IV sets and estimate causal effects accurately. Experimental results demonstrate the effectiveness of the method in handling bi-directional causal relationships, providing insights crucial for improving causal inference in complex systems.

Strengths

The main contribution of the paper is presenting sufficient and necessary conditions for the identifiability of the bi-directional model, enabling both valid IV sets for each direction. They also propose a practical and effective cluster fusion-like algorithm for unbiased estimation based on the theorems and prove the correctness of the algorithm. The paper also validates the theoretical findings using extensive experiments on synthetic data along with comparisons to baseline methods. Overall, the paper is well written and easy to follow.

Weaknesses

The paper has no major weaknesses. However, the setting is restricted with causal relations limited to being linear and assuming genetic variants are randomized, which limits the practical applicability of the proposed approach. Additionally, the experimental results provided are mainly synthetic in nature.

Questions

I have few Questions/Suggestions for the Authors: * In line 106, the authors mention, "Following Hausman [1983], we assume that $\beta_{X ->Y} \beta_{X->X} \neq 1$." It would be useful to discuss in the main paper why this assumption is necessary and what happens when it is violated. * Similarly, regarding Assumption 3, the authors mention it as a very natural condition that one expects to hold for the unique identifiability of valid IVs. It would be useful to explain briefly in the main paper why this assumption is necessary for the identifiability of IVs. * The authors in Section 5 claim that with dependence between genetic variants, main results may still be effective in identifying valid IV sets. Does this claim still hold when there is confounding among the genetic variants or between the genetic variants and some phenotype? Or does the dependence just mean direct causal effect here?. * At the moment, the proposed solution is restricted to linear causal relationships. Can one apply the proposed method using some linearization technique for scenarios where causal relationships are not necessarily linear?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors clearly state all the assumptions. The paper could benefit from adding some more discussion on the necessity of these assumptions in the main paper. I don't think the paper has any potential negative societal impacts.

Reviewer St5u7/10 · confidence 4/52024-07-12

Summary

The authors take up a _very useful_ topic, of trying to identify instruments in models where bidirectional adjacencies exist, at least for the Mendelian randomization application.

Strengths

The topic of the paper is on point--this is something we need to know more about, as bidirectional edges obviously exist in real data. This was an excellent paper, thanks. The discussion made sense to me from start to finish, and the experimental results were compelling. Thanks.

Weaknesses

From _my_ perspective, there were no glaring weaknesses to this paper. Perhaps other reviewers have issues to mention. The only possible weakness I saw was the strong reliance on the assumption of linearity, though in the discussion this was mentioned as an assumption that could possibly be relaxed in future work.

Questions

No particular questions.

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

I did not see a discussion of societal impact.

Reviewer XbcF6/10 · confidence 4/52024-07-13

Summary

The paper addresses the problem of estimating causal effects in bi-directional Mendelian randomization (MR) models with some invalid instrumental variables (IVs) and unmeasured confounding. It proposes a framework for identifying valid IV sets under the assumption that the IV set consists of genetic variants that are independent of each other and that at least two of them are valid IVs. The authors introduce a cluster fusion-like algorithm based on this framework and demonstrate its effectiveness through theoretical proofs and experimental results.

Strengths

The authors establish both necessary and sufficient conditions for identifying bi-directional Mendelian randomization (MR) models, which builds upon previous work focusing on uni-directional MR. The proposed cluster fusion-like algorithm is well-founded. The experimental results on synthetic datasets show the algorithm's efficacy in estimating causal effects. These results support the theoretical claims and suggest that the method performs well in practice.

Weaknesses

While the paper discusses various assumptions (such as the independence of genetic variants and existence of two valid IVs), it would benefit from a more in-depth exploration of the limitations and potential pitfalls of these assumptions in real-world data. Addressing how violations of these assumptions impact the results could strengthen the paper. The experiments are performed on synthetic datasets. While this is a good starting point, additional validation on real-world datasets would provide more robust evidence of the method’s practical utility.

Questions

1) How does your method perform when the assumptions (e.g., independence of genetic variants) are violated in practice? Are there any robust techniques or adjustments to handle such cases? 2) Regarding the construction of the IV set, do you find that a larger IV set generally leads to more robust causal estimates, or does it introduce more complexity and potential for bias with invalid instruments? 3) Is the process of constructing the IV set dependent on the order in which instruments are considered? Specifically, does sequentially adding IVs versus a simultaneous assessment of all potential IVs impact the validity and effectiveness of the identified set? 4) Can the proposed algorithm handle large-scale datasets efficiently? What are the computational complexities and potential bottlenecks?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

See weaknesses.

Reviewer Qa6w7/10 · confidence 4/52024-07-14

Summary

This paper studies the identifiability problem of the bi-directional Mendelian randomization (MR) model, where $X$ and $Y$ are a pair of phenotypes of interest and causes of each other, and $\textbf{G}$ is the set of measured genetic variants, which may include invalid instrumental variables (IVs). Under some assumptions, the paper has identified and proved correct the sufficient and necessary conditions for identifying valid IV sets from $\textbf{G}$ based on observational data, without requiring prior knowledge about which candidate IVs in $\textbf{G}$ are valid or invalid. Supported by the theoretical result, an algorithm is proposed for finding the valid IV sets from the set of measured genetic variants $\textbf{G}$ using observation data and estimating the bi-directional causal effects using the found valid IVs. Experiments are conducted with synthetic data to show the effectiveness of the proposed algorithm.

Strengths

1. The paper addresses a challenging and practical problem. 2. The work is comprehensive, with both theoretical results and corresponding algorithm presented. 3. The paper is very well written in general.

Weaknesses

1. The experimental evaluation is done with synthetic data only. Although the presented experiments with synthetic data are comprehensive and the identification conditions have been theoretically proved, as the theoretical result relies on several assumptions, it would be necessary to conduct some case studies with real world data to evaluate how the method works in practice (where domain knowledge or literature can be used to justify the correctness of the found IV sets) 2. It would be very helpful if the assumptions and their feasibility (and consequences/limitations) in practice can be illustrated and justified with real world examples.

Questions

1. Line 106: Does the assumption regarding the multiplication of the two effects have any practical meaning/implication? 2. Could you explain what "cluster fusion" means exactly in the paper and why the proposed algorithm is said to be "cluster fusion-like"? 3. Section 6.2 - how the one-directional data used in this section generated? 4. The work is based on the assumed structure in Figure 2 (plus some invalid IVs as illustrated in the other figures) , but in practice there would be more complicated situations than those, e.g. the vertical pleiotropy effect in biology where the IVs (genetic variants) are associated with another phenotype (or biological pathway) and this in turn causes the two phenotypes of interest ($X$ and $Y$).

Rating

7

Confidence

4

Soundness

3

Presentation

4

Contribution

3

Limitations

Some limitations of the paper have been discussed briefly, but as mentioned above, the consequence and limitations due to the assumptions should be discussed a bit more.

Reviewer UQEn2024-08-08

Re.

Thanks for responding to my questions. The suggested changes by the authors, including additional experiments and clarifications, would be very beneficial for the paper. I will keep my decision and score for the paper.

Reviewer Qa6w2024-08-12

Thanks for your responses

Thanks the authors for your detailed responses. The extra experiments and discussions will be very helpful. I am happy to keep my positive rating.

Reviewer St5u2024-08-12

Thanks.

Thanks for the rebuttal. I will stick to my original assessment.

Reviewer XbcF2024-08-12

Thank you for the responses.

Thank you for your detailed responses. I appreciate the clarification and will maintain my positive score.

Authorsrebuttal2024-08-13

Thank all reviewers again for greatly improving our paper!

Dear all reviewers, We sincerely appreciate all your positive comments! We are grateful for your valuable and inspiring suggestions, which are of great help in improving our paper! Best wishes, The authors

Program Chairsdecision2024-09-25

Decision

Accept (oral)

© 2026 NYSGPT2525 LLC