Summary
This work proposes a method for adapting stopping tests of randomized experiments in heterogeneous populations. Specifically, the authors motivate the problem, namely why heterogeneous treatment effects lead to late stopping of randomized experiments, for instance when a minority group is harmed. They then propose a two stage method which first predicts a weighting of the original test statistic components used in stopping decisions, then and then uses these statistics to make the stopping decision. The methodological contribution is well-motivated and supported by theoretical results which analyse the convergence behavior of these weights, and the probability of stopping under the assumption of knowing the group membership.
Strengths
* There is a lack of machine learning methods that addresses the heterogeneous early stopping problem, and this paper provides a possible, first solution while making few assumptions. This renders the work an original, well-motivated contribution.
* The paper is very clear and well written. In particular, the links between each of the sections are very clear.
* The experimental results, including the simulated scenarios, are convincing and interesting. For instance Figure 2 makes clear why the proposed approach is advantageous over homogeneous stopping tests.
Weaknesses
* The task is somewhat niche. It is furthermore unclear to what degree the stopping task in randomized experiments could be reformulated as a similar task in another domain, i.e. to what degree this or similar problems have been solved in other contexts.
* Overall, the work makes various idealised assumptions in both the theoretical results, and in the (synthetic) experiments considered. It is an interesting proof of concept, but there is a lot of work which would need to be done to make this method applicable in practical settings, for instance real-world clinical trials which this method is motivated by, but does not evaluate on. Two further points to note is the lack of performance on high-dimensional data that the authors themselves notes, which is present in some clnical trials. Second, the method crucially relies on treatment effect estimation methods in Stage 1, which themselves are far from being widely applicable in practice.
Questions
* The problem setup of CLASH stops the entire experiment if a stopping decision has been made. However, in clinical trials for example, it may not be ethical to stop the experiment for a majority group which benefits from the treatment. Could CLASH be adapted to an online setting where the experiment is only stopped for the harmed subgroup? It would be interesting if the authors could comment on this.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
* Prop. 3.1 assumes the group membership (and CATE) is known, which is never the case in practice. The authors state this, yet it is unclear how one would efficiently infer group membership in large-scale scenarios in practice, and whether the proposed solution in Stage 1 of the algorithm would work. It is consequently unclear how well the proposed tests would generalise if group membership is not known, or how the asymptotic behavior analysed in this Proposition would change in this case. It would be helpful if the authors could comment on this.
* The experimental settings are all limited to two groups. In many real-world settings such as clinical trials, we would expect more than two groups with somewhat homogeneous treatment effect. How this method would perform in such cases is unclear, and neither discussed in the paper.