Summary
This paper studies the problem of “Adaptive Neyman Allocation”, which involves designing an efficient, adaptive experimental design. Neyman allocation is an infeasible experimental design which would be optimal (minimum variance) if the planner knew all the exact potential outcomes under different treatments. However, this is infeasible, and so the goal considered in the paper is to build an adaptive experimental design which is nearly as efficient (in terms of variance) as the infeasible non-adaptive Neyman allocation asymptotically.
To measure the performance, the first contribution of the paper is to propose new measures of regret (and regret ratio) similar to the notion of regret in bandits/statistical learning. Second, the paper proposes an adaptive design based on the idea of adaptive gradience descent to adjust the treatment probabilities to minimize regret. The paper shows that the regret in this approach scales as O(sqrt(T)). Finally the paper constructs confidence intervals which guarantee asymptotic coverage of the average treatment effect.
Strengths
– Novelty: The paper claims to be the first to introduce the notion of Neyman regret in the context of adaptive experimental designs. I am not fully aware of the related literature, but this seems to be an interesting contribution to analyze from this perspective.
– Well-written: The paper is well organized overall and the concepts are introduced and explained crisply.
– Theory: The paper presents results well-grounded in theory.
Weaknesses
1. Related Work: The related work section mainly talks about Neyman allocation and about casual inference under adaptively collected data. However, it seems there is also a large body of work that studies adaptive experimentation/ adaptive design for randomized trials. All of these references seem to be missing (please see few examples below and references therein)? It would be good to distinguish this paper from this body of related work.
-- Eggenberger, Florian, and George Pólya. "Über die statistik verketteter vorgänge." ZAMM‐Journal of Applied Mathematics and Mechanics/Zeitschrift für Angewandte Mathematik und Mechanik 3, no. 4 (1923): 279-289
-- Xu, Yanxun, Lorenzo Trippa, Peter Müller, and Yuan Ji. "Subgroup-based adaptive (SUBA) designs for multi-arm biomarker trials." Statistics in Biosciences 8 (2016): 159-180.
-- Eisele, Jeffrey R. "The doubly adaptive biased coin design for sequential clinical trials." Journal of Statistical Planning and Inference 38, no. 2 (1994): 249-261.
-- Hu, Feifang, and William F. Rosenberger. "Optimality, variability, power: evaluating response-adaptive randomization procedures for treatment comparisons." Journal of the American Statistical Association 98, no. 463 (2003): 671-678.
2. Unsurprising: To me it is a little unsurprising/ unimpressive that the approach can achieve the optimal data efficiency as T tend to infinity (after a really large number of samples). The authors also seem to acknowledge this in part in the final section of the paper.
3. Empirical Evaluation: It may be useful to add more interesting baselines if available in comparison of the regret
Questions
– In section 4.1 why isn’t the variance of adaptive experimental design, V also a function of t?
– There seem to be a few other approaches that study adaptive randomziation in clinical trials. For example: Zhang, Lanju, and William F. Rosenberger. "Response‐adaptive randomization for clinical trials with continuous outcomes." Biometrics 62.2 (2006): 562-569.
How does the proposed approach compare against existing methods?
– Ethical concern: Please see additional question under Limitations section.
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
Could there be fairness/social welfare concerns stemming from such optimal designs? Perhaps the allocation algorithm may assign a higher probability to a less effective or detrimental treatment simply because the variance in its outcomes is higher. As a result, in the quest for minimized variance, could the negative treatment be administered more often than advisable/necessary?