Clip-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments

From clinical development of cancer therapies to investigations into partisan bias, adaptive sequential designs have become increasingly popular method for causal inference, as they offer the possibility of improved precision over their non-adaptive counterparts. However, even in simple settings (e.g. two treatments) the extent to which adaptive designs can improve precision is not sufficiently well understood. In this work, we study the problem of Adaptive Neyman Allocation in a design-based potential outcomes framework, where the experimenter seeks to construct an adaptive design which is nearly as efficient as the optimal (but infeasible) non-adaptive Neyman design, which has access to all potential outcomes. Motivated by connections to online optimization, we propose Neyman Ratio and Neyman Regret as two (equivalent) performance measures of adaptive designs for this problem. We present Clip-OGD, an adaptive design which achieves $\widetilde{O}(\sqrt{T})$ expected Neyman regret and thereby recovers the optimal Neyman variance in large samples. Finally, we construct a conservative variance estimator which facilitates the development of asymptotically valid confidence intervals. To complement our theoretical results, we conduct simulations using data from a microeconomic experiment.

Paper

References (40)

Scroll for more · 28 remaining

Similar papers

Peer review

Reviewer wYBx6/10 · confidence 2/52023-07-02

Summary

This paper studies adaptive Neymann allocation and proposed an algorithm that achieves expected Neymann regret $\tilde{O}(\sqrt{T})$. I am not an expert in sequential experiment design so my ability is limited to assess the impact/relevance of this paper. Yet, I do consider myself well-versed in the potential-outcomes framework.

Strengths

- I found the writing and analyses appear to be rigorous and precise. - The proposed algorithm achieves the Neymann variance in the synthetic experiment.

Weaknesses

See "Questions" for more details.

Questions

1. I understand the authors may wish to provide sufficient background, but it is a bit hard to tell what are existing (and well-known) results and what are new results (except starting from page 6 when the authors introduce the algorithm). 2. I felt Assumption 1 is a bit tricky to decipher: what is the motivation to have the same constant $c$? What are the scenarios where this assumption may break? 3. Have you considered combining CLIP-OGD with an outcome model to get a doubly-robust estimator? 4. Can you add more simulations?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

Not applicable

Reviewer etbf7/10 · confidence 3/52023-07-06

Summary

The authors study adaptive allocation of samples into control and treatment groups in an experiment. They (i) define new performance measures for such allocations, namely Neyman ratio and Neyman regret, (ii) introduce a new algorithm called Clip-OGD that achieves optimal Neyman regret, and (iii) provide asymptotically correct confidence intervals for experiments run with Clip-OGD.

Strengths

Presentation-wise, this is an extremely well-written paper. In particular, (i) The metrics that are proposed, Neyman ratio and Neyman regret, are well motivated. It is convincing why minimizing the Neyman ratio is desirable and how minimizing Neyman regret is a surrogate for that objective. (ii) Clip-OGD is a simple and intuitive algorithm. Notably, it does not have hyper-parameters that require tuning (although they could potentially be tuned). This is a major strength for an online algorithm where cross-validation using an offline dataset would not be possible.

Weaknesses

There are only minor weaknesses (please see my questions below).

Questions

1) In the paragraph starting from line 226, Neyman regret is discussed from the perspective of multi-armed bandits. I understand how explore-exploit algorithms like UCB are not suitable for minimizing the Neyman regret; the purpose of adaptive Neyman allocation seems to be purely exploratory. Then, how does it relate to existing pure-exploration bandits? 2) Assumption 2 is not intuitive and is not explained well enough. What does "ruling out settings where the outcomes were chosen with knowledge of Clip-OGD" mean exactly? I understand the formulation assumes $y_0(t)$ and $y_1(t)$ are deterministic, but just to understand Assumption 2, what would its equivalent have been under a super-population assumption (i.e. if $Y_0(t)$ and $Y_1(t)$ were to be random and i.i.d.)? 3) Confidence intervals provided in Section 5 seem to be only asymptotically correct. Does this mean that no finite-sample guarantees are possible regarding the error rate of these intervals? This would limit the use of Clip-OGD in settings where strict error control is required (such as clinical trials). 4) Why do the experiments not include a multi-armed bandit baseline although they are mentioned as an alternative design in Section 6? 5) What does "informal" mean in Proposition 6.1? If it does not have a formal proof, maybe it should be stated as a remark rather than a proposition.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Yes, the authors adequately addressed the limitations of their work.

Reviewer 2JiC7/10 · confidence 3/52023-07-08

Summary

The work considers adaptive experimental design for sequential experiments. To do so, a new regret-like measure, called Neyman regret is defined that compares the ratio of the variance under the chosen experiment design with respect to the variance under the optimal experiment design. Drawing connections to online convex optimization literature, an algorithm is developed that provides $\tilde O(\sqrt(T))$ Neyman regret. In contrast, it is also established that two-stage explore then commit or multi-arm bandit setups may not result in (super)linear Neyman regret. Additionally, asymptotically valid confidence intervals using the adaptively collected data are provided.

Strengths

S1. An important topic of sequential experimental design, particularly setting up the framework from the point of regret minimization. S2. The core idea is well-presented S3. Comparison of Neyman regret with alternative designs is also valuable.

Weaknesses

W1. Numerical simulations that analyzed more aspects of the algorithm could have made the paper stronger.

Questions

A. For the gradient term in the algorithm, it might be beneficial to elaborate on that, i.e., from variance eqn to eqn for estimating that variance and taking the derivative to get the cubic terms. B. I see the need for the projection term in the algorithm, but it is not clear how to actually set it in practice. C. I am curious if the authors can elaborate on the choice of the utility for gradient estimation. Instead of minimizing the utility with just the new sample, why not do a Follow-the-leader style algorithm to minimize the utility over all the samples seen so far? Maybe even FTRL, with a regularizer that prevents the distribution from shifting too far from Bernoulli(0.5) could replace the projection step? D. I think the discussion around non-superefficient variance is important and I while I see the need of it, I do not quite understand it properly. Having more discussion about that could be beneficial. E. What is the practical relevance of the experiment setup for Fig 1(b)?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

F. To understand the sensitivity of the projection hyper-parameter, can ablations for it be provided in Fig 1?

Reviewer NsFr4/10 · confidence 3/52023-07-12

Summary

This paper studies the problem of “Adaptive Neyman Allocation”, which involves designing an efficient, adaptive experimental design. Neyman allocation is an infeasible experimental design which would be optimal (minimum variance) if the planner knew all the exact potential outcomes under different treatments. However, this is infeasible, and so the goal considered in the paper is to build an adaptive experimental design which is nearly as efficient (in terms of variance) as the infeasible non-adaptive Neyman allocation asymptotically. To measure the performance, the first contribution of the paper is to propose new measures of regret (and regret ratio) similar to the notion of regret in bandits/statistical learning. Second, the paper proposes an adaptive design based on the idea of adaptive gradience descent to adjust the treatment probabilities to minimize regret. The paper shows that the regret in this approach scales as O(sqrt(T)). Finally the paper constructs confidence intervals which guarantee asymptotic coverage of the average treatment effect.

Strengths

– Novelty: The paper claims to be the first to introduce the notion of Neyman regret in the context of adaptive experimental designs. I am not fully aware of the related literature, but this seems to be an interesting contribution to analyze from this perspective. – Well-written: The paper is well organized overall and the concepts are introduced and explained crisply. – Theory: The paper presents results well-grounded in theory.

Weaknesses

1. Related Work: The related work section mainly talks about Neyman allocation and about casual inference under adaptively collected data. However, it seems there is also a large body of work that studies adaptive experimentation/ adaptive design for randomized trials. All of these references seem to be missing (please see few examples below and references therein)? It would be good to distinguish this paper from this body of related work. -- Eggenberger, Florian, and George Pólya. "Über die statistik verketteter vorgänge." ZAMM‐Journal of Applied Mathematics and Mechanics/Zeitschrift für Angewandte Mathematik und Mechanik 3, no. 4 (1923): 279-289 -- Xu, Yanxun, Lorenzo Trippa, Peter Müller, and Yuan Ji. "Subgroup-based adaptive (SUBA) designs for multi-arm biomarker trials." Statistics in Biosciences 8 (2016): 159-180. -- Eisele, Jeffrey R. "The doubly adaptive biased coin design for sequential clinical trials." Journal of Statistical Planning and Inference 38, no. 2 (1994): 249-261. -- Hu, Feifang, and William F. Rosenberger. "Optimality, variability, power: evaluating response-adaptive randomization procedures for treatment comparisons." Journal of the American Statistical Association 98, no. 463 (2003): 671-678. 2. Unsurprising: To me it is a little unsurprising/ unimpressive that the approach can achieve the optimal data efficiency as T tend to infinity (after a really large number of samples). The authors also seem to acknowledge this in part in the final section of the paper. 3. Empirical Evaluation: It may be useful to add more interesting baselines if available in comparison of the regret

Questions

– In section 4.1 why isn’t the variance of adaptive experimental design, V also a function of t? – There seem to be a few other approaches that study adaptive randomziation in clinical trials. For example: Zhang, Lanju, and William F. Rosenberger. "Response‐adaptive randomization for clinical trials with continuous outcomes." Biometrics 62.2 (2006): 562-569. How does the proposed approach compare against existing methods? – Ethical concern: Please see additional question under Limitations section.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

Could there be fairness/social welfare concerns stemming from such optimal designs? Perhaps the allocation algorithm may assign a higher probability to a less effective or detrimental treatment simply because the variance in its outcomes is higher. As a result, in the quest for minimized variance, could the negative treatment be administered more often than advisable/necessary?

Reviewer xwuB5/10 · confidence 1/52023-07-27

Summary

The paper proposes a new adaptative Neyman allocation for experimental design. The proposed adaptative design gets close to the optimal non-adaptative strategy without suffering from the same infeasibilities.

Strengths

1) The writing and structure of the paper are OK. 2) The topic is very relevant, and the paper proposes adaptative variance with formal guarantees ensuring feasibility is interesting.

Weaknesses

1) The paper is not easy to follow for the non-expert, both concerning the notations used and the derivations. 2) It seems that the motivation for this work could be broader. It is unclear to the reviewer why the authors focused on medical applications. 3) The method is only applied to a single microeconomic example.

Questions

1) I would like to see some discussion about the medical applications vs a more broader motivation.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

Limitations are adequately addressed.

Reviewer wYBx2023-08-16

Thank you for the reply

I thank the authors for the detailed response and addressing my concerns. I agree that with a more streamlined and detailed discussions on Assumption 1, and with the new results this paper is stronger that it was.

Authorsrebuttal2023-08-17

response to reply

Thank you for your response. We are happy to hear that you found that the revisions brought forth in the rebuttals increased the strength of the paper.

Area Chair b83D2023-08-21

Hello Reviewer NsFr, Were your concerns regarding the novelty (surprise) addressed by the authors? I have reviewed the related works you suggested in the context of this paper, and while they are relevant, I do not believe their omission warrants rejection alone. Given the authors' efforts to bulk up the related work, are your concerns addressed?

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC