Targeted Sequential Indirect Experiment Design

Scientific hypotheses typically concern specific aspects of complex, imperfectly understood or entirely unknown mechanisms, such as the effect of gene expression levels on phenotypes or how microbial communities influence environmental health. Such queries are inherently causal (rather than purely associational), but in many settings, experiments can not be conducted directly on the target variables of interest, but are indirect. Therefore, they perturb the target variable, but do not remove potential confounding factors. If, additionally, the resulting experimental measurements are multi-dimensional and the studied mechanisms nonlinear, the query of interest is generally not identified. We develop an adaptive strategy to design indirect experiments that optimally inform a targeted query about the ground truth mechanism in terms of sequentially narrowing the gap between an upper and lower bound on the query. While the general formulation consists of a bi-level optimization procedure, we derive an efficiently estimable analytical kernel-based estimator of the bounds for the causal effect, a query of key interest, and demonstrate the efficacy of our approach in confounded, multivariate, nonlinear synthetic settings.

Paper

Similar papers

Peer review

Reviewer sEkT5/10 · confidence 3/52024-07-09

Summary

This paper designs comprehensive experiments that maximize the information gained about the query of interest within a fixed budget of experimentation, including nonlinear, multi-variate, confounded settings.

Strengths

This paper is well-motivated and relevant to the causal inference. Overall, this work is well-written and organized.

Weaknesses

1. Are the theoretical results practically useful when addressing real-world problems? 2. How does the proposed method handle non-convex optimization, and what are the assumptions on the loss function? 3. Can the authors provide more intuitive theoretical explanations of the difference between convex and non-convex optimization in their settings? 4. How can the effectiveness of the proposed method be guaranteed when the dimension p is significantly larger than the number of training samples?

Questions

1. Are there any crossing problems when minimizing the optimization problem in Equation (5)? How do we ensure that $Q^{+}(\pi)$ is always larger than $Q^{-}(\pi)$? 2. Theoretically, are there any requirements for the dimensions of $Z$ and $U$? In your experiments, you only consider a low-dimensional setting. Can you add any experiments for high dimensions for $Z$ and $U$? 3. For Theorem 2, the proposed method involves several tuning parameters, such as $λ_g,λ_f,λ_c$. This step can be quite worrying for the practical implementation of the proposed method. 1) it can be time-consuming and 2) setting the candidate values is likely quite subjective. The authors should conduct more thorough studies and provide more guidelines on how the choice of tuning parameters affects the method's performance. For example, the authors have selected $λ_g=0.1,λ_f=0.1,λ_c=0.1$, and the learning rate $α_t=0.01$. This selection process appears too subjective. 4. The experiments are not sufficient for several reasons: (1) The number of replications used in the experiments is quite small (only 25), which raises concerns about the time efficiency of the proposed method. (2) The paper does not include a real-world data analysis to validate the effectiveness of the proposed algorithm from a practical perspective.

Rating

5

Confidence

3

Soundness

2

Presentation

3

Contribution

3

Limitations

1. Theoretical requirements for the dimensions of $Z$ and $U$ are not discussed. Experiments only consider a low-dimensional setting, see questions 2. The proposed method involves several tuning parameters, which can be problematic for practical implementation due to time consumption and the subjective setting of candidate values; see questions. 3. Insufficient Experiments: see questions.

Authorsrebuttal2024-08-07

4 **Insufficient experiments:** Our runtime evaluation in appendix C, Fig. 4 (see also Fig. 3 in the rebuttal pdf), shows that the computational cost of our approach is not a severe limitation, i.e., runtime is not a real concern. While 25 replications already provide a good assessment of the finite sample variance, we happily increase this number to 100 for the revised version. From the first finished settings, the results are visually unchanged. We agree that we would have loved to include a real-world application of the method, but point to our discussion of the fundamental difficulty of assessing our method in a real-world setting in our “limitations” setting (l.392 and following).

Reviewer sEkT2024-08-14

Thanks for the author's response, which addressed most of my concerns. I will maintain my score.

Reviewer aNmJ6/10 · confidence 3/52024-07-12

Summary

This paper proposes a framework for designing sequential indirect experiments for estimating targeted scientific queries in complex, nonlinear environments with potential unobserved confounding when direct intervention is impractical or impossible. The authors formulate the problem as the sequential instrument design, using instrumental variables and minimax optimization to estimate the upper and lower bounds of targeted causal effects. The proposed method then tightens these bounds iteratively through adaptive strategies. Experiments on simulated data demonstrate the proposed method's efficacy compared to non-adaptive experimental design baselines.

Strengths

1. This paper focuses on targeted sequential indirect experimental design, which is an important problem in scientific discovery where direct intervention is often impractical or impossible. 2. The proposed method is more flexible as it considers nonlinear, multi-variate, and confounded settings, and formulating the problem as a sequential underspecified instrumental variable estimation is intuitive and sound. 3. The authors develop closed-form estimators for the bounds given targeted queries when the mechanism is in an RKHS. 4. The proposed method outperforms non-adaptive baselines in synthetic experiments.

Weaknesses

1. Although the proposed method focuses on indirect experimental design, it still requires the causal structure between variables known, which is also one of the main challenges in scientific discovery. I am curious about the method's sensitivity to imperfections in the causal structure, assumptions regarding instrumental variables, and the presence of confounders. It would be great if the authors could discuss further about them. 2. The authors only conduct synthetic experiments in a simple setting when the number of variables is small. I wonder if the authors could conduct more experiments in a more complex setting with a larger number of nodes, and also discuss the scalability of the proposed method (since kernel-based techniques are used for optimization as mentioned in the paper).

Questions

1. Please see the questions in the Weaknesses part. 2. It seems that the proposed method is closely related to causal Bayesian Optimization [1]. I wonder if the authors could discuss the connection in detail. 3. Experiments demonstrate that the proposed method performs well for local causal effect queries. However, it is usually unknown whether the target causal effect is local or more global (i.e., long-range) in real-world applications. I wonder if the proposed method could estimate long-range causal effects accurately. 4. What is the difference/connection between the proposed method and targeted indirect experiment design in an active learning setting? [1] Aglietti, V., Lu, X., Paleyes, A., & González, J. (2020, June). Causal bayesian optimization. In International Conference on Artificial Intelligence and Statistics (pp. 3155-3164). PMLR.

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors adequately addressed the limitations of their work.

Authorsrebuttal2024-08-07

[1] Cinelli, Carlos, and Chad Hazlett. "An omitted variable bias framework for sensitivity analysis of instrumental variables." Available at SSRN 4217915 (2022). [2] Vancak, Valentin, and Arvid Sjölander. "Sensitivity analysis of G‐estimators to invalid instrumental variables." Statistics in Medicine 42.23 (2023): 4257-4281. [3] Kilbertus, Niki, Matt J. Kusner, and Ricardo Silva. "A class of algorithms for general instrumental variable models." Advances in Neural Information Processing Systems 33 (2020) [4] Drineas, P., Mahoney, M.W. (2005). On the Nystrom Method for Approximating a Gram Matrix for Improved Kernel-Based Learning, 2005, Journal of Machine Learning Research, http://jmlr.org/papers/v6/drineas05a.html [5] Li, Mu et al. “Large-Scale Nyström Kernel Matrix Approximation Using Randomized SVD.” IEEE Transactions on Neural Networks and Learning Systems 26 (2015): 152-164. [6] Elisabeth Ailer, Jason Hartford, and Niki Kilbertus. 2023. Sequential underspecified instrument selection for cause-effect estimation. In Proceedings of the 40th International Conference on Machine Learning (ICML'23), Vol. 202. JMLR.org, Article 19, 408–420.

Reviewer pD3c7/10 · confidence 3/52024-07-12

Summary

The authors provide a procedure to use a sequence of encouragement designs to identify target functionals about a particular causal relationship.

Strengths

This is a really cool problem setting. It isn't obvious to me that it's particularly common, but I think the larger idea of trying to think about _which_ experiments to run in order to gain knowledge about the world is a worthy line of inquiry. In general, I think the approach of getting partial identification on a parameter and then determining a series of actions to take in order to reduce the width of that bound is a fantastic approach. This is a great way to conceptualize a variety of problems, I suspect.

Weaknesses

It isn't clear to me why Equation 4 is a bound on Q[f_0]. Maybe this should be obvious to me, but I think if it isn't obvious to me, it probably won't be obvious to a lot of your readers, as I think I've read more into this literature than most people. This feels like your big contribution: after this bound is setup, the remainder isn't what I would call straightforward, but it feels more like standard machinery that I'm used to. Conditional on the bound, Theorem 1, 2 and the experiment selection procedures you provide all make a lot of sense. But I just don't see where this bound comes from. I would like to see (i) a clearer explanation of why this expression bounds Q[f_0], (ii) some clearer examples of the cases in which the bounds are equal and Q[f_0] is identified. I think it might also help to lay out when a non-optimal policy (what you call "non-informative experimentation") doesn't identify Q[f_0]. This is a bit confusing to me, as I'd expect that complete randomization as a policy would provide identification so long as there is positivity across the whole space: one could do something like off-policy evaluation to identify any particular \pi(Z). I believe the problem comes with the restriction discussed at the start of Section 3.2, but I do not see why it is that the solution to ensure that r_0 lies in TT^* is to take the +\- of Q[f] in the objective as from Eq 3. There's a connection here that you have not clearly spelled out to me. Maybe its possible that the problem is that I need to read Bennett et al 2023 much more carefully to see why this follows, but that isn't a fair ask for your readers who are reading _this_ paper rather than that one. Unfortunately, it's quite difficult for me to evaluate the overall novelty of this work without understanding this component. I think you have something very cool here, but I can't quite work that out based on the paper as it stands. I'd like to reiterate that I follow what you're doing at a high level (i.e. line 96-106 makes sense in general), but I don't follow when you get into specifics. Some other miscellaneous thoughts that aren't as important: - Do you _have_ to call this (\partial_i f)(x^*) a local effect? Between LATEs and everything else, this feels like an absurdly overloaded term. - The experiments section makes it difficult to see clear quantitative comparisons between methods. I would like to see things like causal mean-squared error. I recognize that's a bit difficult when you're doing partial-ID, but demonstrating that the bounds collapse to the correct Q[f_0] is important. Even just doing something like showing the midpoint of the bounds quantitatively (with confidence interval based on replicates) and the width of the bound would be useful to make sure the process is behaving sensibly.

Questions

see above

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

see above

Authorsrebuttal2024-08-12

We thank the reviewer for their comment and increasing their score. The intuition about the bounds is very well put. We highly appreciate the suggestion and will include the description of the CRISPR use-case into the revised version.

Reviewer GhAo7/10 · confidence 3/52024-07-30

Summary

The authors' primary goal is to design experiments that maximally inform a query of interest about the underlying causal mechanism, within a fixed experimentation budget. They address this by maintaining upper and lower bounds on the query and sequentially selecting experiments to minimize the gap between these bounds. They show that by treating experiments as instrumental variables, they can estimate these bounds using existing techniques in nonlinear instrumental variable estimation. Their procedure involves a bi-level optimization: an inner optimization estimates the bounds, while an outer optimization seeks to minimize the gap between them. For certain queries, when assuming the underlying function lies within a reproducing kernel Hilbert space (RKHS), they derive closed-form solutions for the inner estimation problem. The authors develop adaptive strategies for the outer optimization to iteratively tighten these bounds and demonstrate empirically that their method robustly selects informative experiments, leading to identification of the query of interest when possible within the allowed experimentation framework.

Strengths

S1. The presentation is generally clear, with a logical structure that guides readers through the problem formulation, methodology, and initial experimental results. S2. The authors present a novel approach to causal effect estimation in nonlinear systems with potential unobserved confounding. This framework addresses some challenges in scientific settings where direct experimentation is not possible, potentially aiding more targeted scientific discovery.

Weaknesses

W1. The paper only demonstrates results on a low-dimensional synthetic setting (d_x = d_z = 2). While this provides a proof-of-concept, it's unclear how the method scales to higher dimensions or performs on real-world data. The authors could: - Extend experiments to higher dimensions (e.g. d_x, d_z = 10 or 20) to show scalability. - Test on semi-synthetic data by using real covariates with simulated outcomes. - Discuss computational complexity as dimensionality increases. W2. The paper makes strong assumptions (e.g. RKHS, valid IVs) without thoroughly discussing their implications. The authors could: - Provide sensitivity analyses for key assumptions. - Discuss scenarios where assumptions may not hold in practice. - Clarify which parts of the method rely on which assumptions.

Questions

No.

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

No.

Authorsrebuttal2024-08-07

[1] Drineas, P., Mahoney, M.W. (2005). On the Nystrom Method for Approximating a Gram Matrix for Improved Kernel-Based Learning, 2005, Journal of Machine Learning Research, http://jmlr.org/papers/v6/drineas05a.html [2] Li, Mu et al. “Large-Scale Nyström Kernel Matrix Approximation Using Randomized SVD.” IEEE Transactions on Neural Networks and Learning Systems 26 (2015): 152-164. [3] Bennett, Andrew, et al. "Minimax Instrumental Variable Regression and $ L_2 $ Convergence Guarantees without Identification or Closedness." The Thirty Sixth Annual Conference on Learning Theory. PMLR, 2023. [4] Cinelli, Carlos, and Chad Hazlett. "An omitted variable bias framework for sensitivity analysis of instrumental variables." Available at SSRN 4217915 (2022). [5] Vancak, Valentin, and Arvid Sjölander. "Sensitivity analysis of G‐estimators to invalid instrumental variables." Statistics in Medicine 42.23 (2023): 4257-4281. [6] Kilbertus, Niki, Matt J. Kusner, and Ricardo Silva. "A class of algorithms for general instrumental variable models." Advances in Neural Information Processing Systems 33 (2020) [7] Sriperumbudur, Bharath K., Kenji Fukumizu, and Gert RG Lanckriet. "Universality, Characteristic Kernels and RKHS Embedding of Measures." Journal of Machine Learning Research 12.7 (2011). [8] Schölkopf, Bernhard, and Alexander J. Smola. Learning with kernels: support vector machines, regularization, optimization, and beyond. MIT press, 2002.

Authorsrebuttal2024-08-14

We thank the reviewer for the comment and increasing their score. The connection to Mendelian Randomization can be drawn via the Instrumental Variable setup we use as experiments. MR is a special case of instrumental variables in which genetic variants are used as instruments. Therefore---in theory---our method would propose the next genetic variant which would inform the estimator the most. In practice, MR relies on existing genetic variation within a population and not so much on designing genetic variation. The difference therefore lies in the assumption of the instrument: in our setting we are able to adjust the instrument/experiment, MR looks at existing genetic variation. We will add this aspect in the manuscript as both approaches target similar questions. Thank you for the comment.

Reviewer pD3c2024-08-11

I thank the authors for their clarifications, I think I see more clearly where the bound is coming from. I'll raise my score given this. To put it a bit more plain (to me), you're taking the minimum and maximum norm solutions (by changing the sign of Q[f] in Eq.4) from within the satisfying set of functions $f$ and $g$. Under full identification, those two would optimands coincide, as there is only one satisfying function, but under partial identification that may not be the case. I also think the comment the authors provided to reviewer GhAo regarding the CRISPR use-case is interesting. Coming from more of a social science perspective, this is not a setting which would come up much, but I suspect this may be different in the case you're considering. Making this more clear in the text why this setting matters would help the paper's impact.

Reviewer GhAo2024-08-13

Thank you for your detailed explanation, which has addressed my concerns. I ask this question because recent research in social sciences has highlighted that the validity of instrumental variables remains a significant challenge [1]. I have increased my score accordingly. Also, CRISPR example piqued my interest. I'm curious about how your work relates to or differs from research in biostatistics centered on Mendelian randomization (if any). [1] Rain, rain, go away: 194 potential exclusion-restriction violations for studies using weather as an instrumental variable. Jonathan Mellon.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC