Summary
This paper investigates adaptive experimental design with a focus on fairness. In this context, fairness refers to ensuring that the probabilities of treatment assignment do not significantly differ across various groups. The authors propose a fair adaptive experimental design that simultaneously enhances data use efficiency, achieves an “envy-free” treatment assignment guarantee, and improves the overall welfare of participants.
Strengths
1. Fairness is gaining increasing importance in various fields, including experimental design.
2. The technical parts (Section 2 and Section 3) are solid, well organized and clearly presented.
3. Both simulation studies and a case study based on real data are provided.
Weaknesses
1. Regarding the definition of 'fairness,' this paper considers fairness as the requirement for treatment adoption probabilities to be similar across different groups. However, in many scenarios, particularly in clinical trials, such a definition of fairness may not be entirely convincing. The reason why we want to consider the covariate $X_{it}$ and group $\mathcal{S}_j$ is because the treatment effect may be very different among different groups. For instance, in the case of patients grouped by biomarkers into $\mathcal{S}_1$ and $\mathcal{S}_2$, the new treatment yields a highly positive effect in $\mathcal{S}_1$ but a strongly negative effect in $\mathcal{S}_2$. Consequently, enforcing close treatment probabilities in these two groups might be deemed unfair. It would be beneficial if the authors could provide motivating examples in the introduction that align well with the current formulation.
2. Regarding the consideration of 'welfare,' I have some uncertainty about its direct relevance to the main topic of 'fairness.' While I understand that enforcing fairness constraints may potentially impact welfare, I'm unsure about the specific implications if we were to remove the welfare constraint in Problem A. In lines 66-70, the authors briefly touch upon this question, but the points made are not entirely clear to me. It would be beneficial if the authors could further elaborate on why “the second fairness concern arises when the adaptive treatment allocation does not adequately account for the overall welfare of experimental participants.”
3. About the feasibility of the optimization Problem A and Problem B. In my view, it seems that Problem A could be infeasible, especially when $c_1$ is very small or $c_2$ is very close to ½. I am curious about whether the authors once met the infeasibility issue in the simulation studies and the case study.
4. Regarding the length of the first stage, it would be beneficial to include additional comments in the paper on how to select the value of $n_1$. Furthermore, I would appreciate a clearer explanation of how exactly $n_1$ influences the main results, as this aspect is important for my understanding. Additionally, in lines 107-108, the authors mention that 'An important methodological and practical innovation of our framework is that it does not require the number of participants enrolled in the first stage to be proportional to the overall sample size.' However, more supporting comments on this claim within the main text are expected. Providing some additional explanation or evidence for this statement would strengthen the paper's argument.
5. Regarding the objectives of the experiment, in lines 76-78, the authors state that 'we propose a fair adaptive experimental design strategy that balances competing objectives: improving fairness, enhancing overall welfare, and gaining efficiency. Our strategy strikes a delicate balance between these trade-offs.' However, it remains unclear what trade-off means among these three objectives. While I can comprehend the trade-off between enhancing overall welfare and gaining efficiency, as it has been previously investigated in [1][2]. I find it unlikely that improving fairness could simultaneously trade off with two objectives that already have a trade-off between them. I believe it would be highly beneficial if the authors could provide further insights into the nature and rationale behind the trade-off that exists among these objectives. Expanding on this aspect would provide a deeper understanding of the decision-making process and the underlying considerations in the proposed strategy.
Reference:
[1] Erraqabi, A., Lazaric, A., Valko, M., Brunskill, E., & Liu, Y. E. (2017, April). Trading off rewards and errors in multi-armed bandits. In Artificial Intelligence and Statistics (pp. 709-717). PMLR.
[2] Simchi-Levi, D., & Wang, C. (2023, April). Multi-armed bandit experimental design: Online decision-making and adaptive inference. In International Conference on Artificial Intelligence and Statistics (pp. 3086-3097). PMLR.
Questions
Apart from the major points above, I have several minor points.
1) The paper is well organized in general. However, I think the writing could be further improved. For the technical parts (Sections 2 and 3), after each lemmas and theorems, it will be better to provide more insights and discussions.
2) In Lemma 1, it is more appropriate to denote the oracle treatment assignment probablity for group j under the proposed welfare constraint as $e^∗ _{j,proposed}$ instead of $e^∗ _{j,alternative}$.
3) For a conference paper, I believe the abstract could be more concise as it is currently a bit long.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.