Adaptive Experimentation When You Can't Experiment

This paper introduces the \emph{confounded pure exploration transductive linear bandit} (\texttt{CPET-LB}) problem. As a motivating example, often online services cannot directly assign users to specific control or treatment experiences either for business or practical reasons. In these settings, naively comparing treatment and control groups that may result from self-selection can lead to biased estimates of underlying treatment effects. Instead, online services can employ a properly randomized encouragement that incentivizes users toward a specific treatment. Our methodology provides online services with an adaptive experimental design approach for learning the best-performing treatment for such \textit{encouragement designs}. We consider a more general underlying model captured by a linear structural equation and formulate pure exploration linear bandits in this setting. Though pure exploration has been extensively studied in standard adaptive experimental design settings, we believe this is the first work considering a setting where noise is confounded. Elimination-style algorithms using experimental design methods in combination with a novel finite-time confidence interval on an instrumental variable style estimator are presented with sample complexity upper bounds nearly matching a minimax lower bound. Finally, experiments are conducted that demonstrate the efficacy of our approach.

Paper

Similar papers

Peer review

Reviewer byyL6/10 · confidence 2/52024-07-11

Summary

This article studies the pure exploration transductive linear bandit problem in the presence of instrumental variables. The authors assume a linear structural equation model on the instrument, treatment and outcome. The proposed method attempts to estimate the parameters in the structural model while optimally designing to learn the best arm (i.e. treatment). The paper is mainly theoretical, but some simulations are provided to demonstrate the method.

Strengths

1. The combination of instrument variables and pure exploration appears novel. 2. The proposed method is guided by finite-time confidence bounds for two-stage least squares. The lower and upper bounds of sample complexity are provided in the paper. 3. The method outperforms the standard methods UCB-OLS and UCB-IV.

Weaknesses

1. Section 3 is very dense. The authors can consider shortening the content before Section 2.2 and use more room to explain Section 3. 2. No real-data demonstration of the proposed method.

Questions

N.A.

Rating

6

Confidence

2

Soundness

3

Presentation

2

Contribution

3

Limitations

N.A.

Reviewer 4vBb6/10 · confidence 4/52024-07-11

Summary

The paper introduces the confounded pure exploration transductive linear bandit (CPET-LB) problem, which addresses the challenges of conducting adaptive experimentation in environments where direct randomization is not possible. From my understanding, the paper studies the best arm identification problem under linear structural model with non-compliance. The paper proposes the algorithms with nearly optimal sample complexity guarantee.

Strengths

1. The paper provides a thorough theoretical analysis, including proofs of finite-time confidence intervals for the estimators and sample complexity bounds. The theoretical contributions are significant and appear to be solid. 2. The problem is practical-relevant and important.

Weaknesses

1. The connections with two streams of the literature should be more clearly stated. The basic problem set up is very classical econometric setting where the non-compliance exists. The analysis framework and some of the tools are very standard in transductive linear bandit problem. 2. The presentation of Section 1 is not easy to follow. I feel the authors try to manage the terminology from causal inference and pure exploration. For example, in Line 78, “measurement” and ”evaluation” are new terminology of the paper and a bit hard to connect with the example proposed in introduction. 3. The authors might want to reconsider the title. The title, from my perspective, is a bit confusing and misleading.

Questions

1. Could you please elaborate a bit more on why the confidence interval in Section 2.2 is novel? I know the traditional 2SLS always using asymptotic normality to construct CI. Is non-asymptotic or asymptotic the key difference?

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

See above.

Reviewer oap77/10 · confidence 4/52024-07-13

Summary

This paper addresses the issue of adaptive experimentation using "encouragement" rather than "compulsion instruction," a scenario commonly encountered in industrial applications. The proposed solution integrates linear bandit algorithms with instrumental variables regressions. The authors provide rigorous theoretical guarantees for their method and demonstrate its superior performance compared to traditional A/B testing and conventional linear bandit algorithms.

Strengths

The scenarios examined by the authors are well-motivated. In practice, numerous situations exist where only encouragement can be employed to influence user decisions. Considering the shift in industry from traditional A/B testing to adaptive experimentation, the study presented here has the potential for significant industry impact. In addition, the proposed algorithm is intuitive, and the theoretical guarantees provided are robust.

Weaknesses

The authors could enhance their study by conducting additional experiments to verify the robustness of their results. Furthermore, there should be a more detailed discussion of the p-values and confidence intervals to strengthen the statistical analysis.

Questions

How to think of p-values/confidence intervals as it is important in an experimentation setting.

Rating

7

Confidence

4

Soundness

4

Presentation

3

Contribution

3

Limitations

The authors adequately addressed the limitations.

Reviewer Y9pp5/10 · confidence 3/52024-07-16

Summary

This paper addresses the problem of pure exploration bandits in the setting of encouragement designs. The authors describe the problem in terms of online instrumental variable regression. Toward this end, the authors derive a finite time confidence interval. Using this as the main tool, the authors then describe a pure exploration transductive bandit. Algorithms are provided for optimizing the pure exploration problem. A number of theoretical results are provided describing the entailed sample complexity of the proposed approach. Empirical results are provided which show strong performance against other baselines.

Strengths

This is a problem that has practical relevance in both industrial and social scientific settings. The task, to my knowledge, is novel in its formulation, and the authors do a nice job of delineating this work from prior art. The authors do a nice job of presenting algorithms and analysis in both the settings of a known and unknown structural models. Further, there is thorough analysis of each of the proposed procedures' properties.

Weaknesses

My concerns are largely around two things: 1. There is a fairly limited experimental evaluation here. It would be helpful if the authors provided a more thorough evaluation of the proposed approaches' behavior across a wider range of settings. 2. The text is a little meandering at times and as a result was a little hard to follow on first read. I would suggest a round of editing in order to improve the narrative and organization of the paper.

Questions

Throughout an E-optimal design is also employed. My main question is whether the choice of E-optimality is done as a matter of convenience, or if it is motivated by aspects of the problem that would favor this criterion over other choices.

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

See above.

Reviewer 4vBb2024-08-11

I have read the rebuttal carefully. Thanks for the clarification. I really appreciate your efforts!

Reviewer oap72024-08-12

Thank you for addressing my questions on the confidence interval. I will keep my score unchanged.

Authorsrebuttal2024-08-13

Dear reviewer Y9pp, As we approach the end of the rebuttal period, we sincerely appreciate this last opportunity to further engage with the reviewer and clarify any outstanding questions or concerns. We thank the reviewer again for acknowledging the novelty of our problem formulation and the thorough analysis throughout. If our rebuttal has adequately addressed the points raised, we would kindly request that the reviewer re-evaluate the score accordingly. We remain open to any additional feedback.

Reviewer Y9pp2024-08-14

Thank you for your answers and clarifications to my questions/concerns. I will leave my score unchanged, but appreciate the points raised by the authors.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC