Fair Adaptive Experiments

Randomized experiments have been the gold standard for assessing the effectiveness of a treatment or policy. The classical complete randomization approach assigns treatments based on a prespecified probability and may lead to inefficient use of data. Adaptive experiments improve upon complete randomization by sequentially learning and updating treatment assignment probabilities. However, their application can also raise fairness and equity concerns, as assignment probabilities may vary drastically across groups of participants. Furthermore, when treatment is expected to be extremely beneficial to certain groups of participants, it is more appropriate to expose many of these participants to favorable treatment. In response to these challenges, we propose a fair adaptive experiment strategy that simultaneously enhances data use efficiency, achieves an envy-free treatment assignment guarantee, and improves the overall welfare of participants. An important feature of our proposed strategy is that we do not impose parametric modeling assumptions on the outcome variables, making it more versatile and applicable to a wider array of applications. Through our theoretical investigation, we characterize the convergence rate of the estimated treatment effects and the associated standard deviations at the group level and further prove that our adaptive treatment assignment algorithm, despite not having a closed-form expression, approaches the optimal allocation rule asymptotically. Our proof strategy takes into account the fact that the allocation decisions in our design depend on sequentially accumulated data, which poses a significant challenge in characterizing the properties and conducting statistical inference of our method. We further provide simulation evidence to showcase the performance of our fair adaptive experiment strategy.

Paper

References (70)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer gk4J6/10 · confidence 4/52023-06-20

Summary

This paper investigates adaptive experimental design with a focus on fairness. In this context, fairness refers to ensuring that the probabilities of treatment assignment do not significantly differ across various groups. The authors propose a fair adaptive experimental design that simultaneously enhances data use efficiency, achieves an “envy-free” treatment assignment guarantee, and improves the overall welfare of participants.

Strengths

1. Fairness is gaining increasing importance in various fields, including experimental design. 2. The technical parts (Section 2 and Section 3) are solid, well organized and clearly presented. 3. Both simulation studies and a case study based on real data are provided.

Weaknesses

1. Regarding the definition of 'fairness,' this paper considers fairness as the requirement for treatment adoption probabilities to be similar across different groups. However, in many scenarios, particularly in clinical trials, such a definition of fairness may not be entirely convincing. The reason why we want to consider the covariate $X_{it}$ and group $\mathcal{S}_j$ is because the treatment effect may be very different among different groups. For instance, in the case of patients grouped by biomarkers into $\mathcal{S}_1$ and $\mathcal{S}_2$, the new treatment yields a highly positive effect in $\mathcal{S}_1$ but a strongly negative effect in $\mathcal{S}_2$. Consequently, enforcing close treatment probabilities in these two groups might be deemed unfair. It would be beneficial if the authors could provide motivating examples in the introduction that align well with the current formulation. 2. Regarding the consideration of 'welfare,' I have some uncertainty about its direct relevance to the main topic of 'fairness.' While I understand that enforcing fairness constraints may potentially impact welfare, I'm unsure about the specific implications if we were to remove the welfare constraint in Problem A. In lines 66-70, the authors briefly touch upon this question, but the points made are not entirely clear to me. It would be beneficial if the authors could further elaborate on why “the second fairness concern arises when the adaptive treatment allocation does not adequately account for the overall welfare of experimental participants.” 3. About the feasibility of the optimization Problem A and Problem B. In my view, it seems that Problem A could be infeasible, especially when $c_1$ is very small or $c_2$ is very close to ½. I am curious about whether the authors once met the infeasibility issue in the simulation studies and the case study. 4. Regarding the length of the first stage, it would be beneficial to include additional comments in the paper on how to select the value of $n_1$. Furthermore, I would appreciate a clearer explanation of how exactly $n_1$ influences the main results, as this aspect is important for my understanding. Additionally, in lines 107-108, the authors mention that 'An important methodological and practical innovation of our framework is that it does not require the number of participants enrolled in the first stage to be proportional to the overall sample size.' However, more supporting comments on this claim within the main text are expected. Providing some additional explanation or evidence for this statement would strengthen the paper's argument. 5. Regarding the objectives of the experiment, in lines 76-78, the authors state that 'we propose a fair adaptive experimental design strategy that balances competing objectives: improving fairness, enhancing overall welfare, and gaining efficiency. Our strategy strikes a delicate balance between these trade-offs.' However, it remains unclear what trade-off means among these three objectives. While I can comprehend the trade-off between enhancing overall welfare and gaining efficiency, as it has been previously investigated in [1][2]. I find it unlikely that improving fairness could simultaneously trade off with two objectives that already have a trade-off between them. I believe it would be highly beneficial if the authors could provide further insights into the nature and rationale behind the trade-off that exists among these objectives. Expanding on this aspect would provide a deeper understanding of the decision-making process and the underlying considerations in the proposed strategy. Reference: [1] Erraqabi, A., Lazaric, A., Valko, M., Brunskill, E., & Liu, Y. E. (2017, April). Trading off rewards and errors in multi-armed bandits. In Artificial Intelligence and Statistics (pp. 709-717). PMLR. [2] Simchi-Levi, D., & Wang, C. (2023, April). Multi-armed bandit experimental design: Online decision-making and adaptive inference. In International Conference on Artificial Intelligence and Statistics (pp. 3086-3097). PMLR.

Questions

Apart from the major points above, I have several minor points. 1) The paper is well organized in general. However, I think the writing could be further improved. For the technical parts (Sections 2 and 3), after each lemmas and theorems, it will be better to provide more insights and discussions. 2) In Lemma 1, it is more appropriate to denote the oracle treatment assignment probablity for group j under the proposed welfare constraint as $e^∗ _{j,proposed}$ instead of $e^∗ _{j,alternative}$. 3) For a conference paper, I believe the abstract could be more concise as it is currently a bit long.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

See previous comments.

Reviewer 62S12/10 · confidence 3/52023-06-28

Summary

The authors delve into an examination of adaptive experimental design subject to the constraints of fairness. They portray a scenario where a researcher, within a sequential experimental context, adapts the allocation between control and treatment groups among a cohort of subjects to optimize the statistical efficiency of the average treatment effect estimation. At the same time, the researcher strives to maintain a balanced proportion of treatments across various demographic groups and maximize the total welfare of these treatments, a condition set forth to uphold fairness. The main contribution of this manuscript lies in the formulation of a novel strategy for treatment allocation that simultaneously adheres to the constraints of fairness. The authors provide a theoretical analysis of their devised method, validating that the average treatment effect estimated from their approach exhibits asymptotic consistency and adheres to a central limit theorem-like convergence toward a particular Gaussian distribution. Empirical studies, leveraging synthetic simulations, demonstrate their method's competitive statistical efficiency against existing approaches while ensuring fairness is upheld. Furthermore, the efficacy of the developed method is also confirmed via empirical simulations that mimic real data, further solidifying the reliability and practicality of their method.

Strengths

This paper investigates the fairness problem in the adaptive experimental design. This problem is well-motivated and of particular interest to NeurIPS's community. The experimental results adequately demonstrate the proposed method's competitive efficiency and guarantee of fairness.

Weaknesses

This paper lacks consideration of essential existing contributions, specifically those of - V. Hadad et al. Confidence Intervals for Policy Evaluation in Adaptive Experiments. Proceedings of the National Academy of Sciences, 2021. - D. Simchi-Levi et al. Multi-armed Bandit Experimental Design: Online Decision-making and Adaptive Inference. AISTATS, 2023. - M. Kato et al. Adaptive Experimental Design for Efficient Treatment Effect Estimation. NeurIPS 2020 Workshop on Causal Discovery & Causality-Inspired Machine Learning. Notably, the findings put forth by Hadad et al. warrant particular attention, given their methodological relevance to this research. Hadad et al. introduce an off-policy estimation of the average treatment effect, a method that works regardless of the treatment assignment mechanism. They also validated the asymptotic consistency and normality of their approach. While Hadad et al.'s method operates without covariates, unlike the scenario in the present paper, this discrepancy does not diminish the applicability of their method. This paper's research problem can be recast into a non-covariate one by transforming the experimental process into group-wise processes. Furthermore, with an appropriately selected set of parameters, the method discussed by Hadad et al. can be reduced to this paper's proposed solution. Consequently, the asymptotic consistency and normality results offered by Hadad et al. would translate to the circumstances of this paper, which seamlessly results in Theorems 1 and 2. Additional insightful points can be gleaned from the work of Kato et al., who propose a treatment assignment mechanism to minimize the variance of the average treatment effect. While unconstrained, their method presents intriguing similarities with the constrained approach of the current paper. Appendix F of Kato et al., particularly noteworthy, introduces an ethical consideration in the treatment assignment process by incorporating a fairness constraint. The omission of these existing methods from the current discussion and the direct applicability of these approaches to the circumstances of this paper leads me to question the originality of the proposed research. The paper would benefit from a comparative evaluation with these significant contributions to ensure that it advances the related field. The present paper is unfortunately beset by a number of quality issues that necessitate substantial revisions. To begin with, the paper neglects to outline the assumptions required for causal inference. Particularly, the authors' methodology employs the inverse propensity score weighting technique. This approach necessitates the fulfillment of certain assumptions, encompassing consistency, no unmeasured confounders, and positivity. The lack of explaining these assumptions calls into question the validity of the resultant findings and negatively impacts the overall quality of the paper. Moving to Section 2.2, it is observed that quality concerns exist. Specifically, Lemma 1's statement introduces the concept of an alternative constraint without providing its rigorous definition. This omission makes us hard to ascertain the correctness of the lemma. Additionally, the section contains considerable redundancy, particularly evident between the contents of Lines 209-219 and Lines 220-230. Such repetition impairs the overall readability of the paper and dilutes the impact of the information conveyed. Cumulatively, these quality issues have contributed to a lower assessment of the paper's presentation score in this review.

Questions

I suggest that the authors elucidate a specific scenario wherein fairness issues arise and illustrate how the applied fairness constraint alleviates these problems. The fairness constraints utilized by the authors appear to diverge from conventional ones; thus, it is essential for the authors to rationalize their choice to implement the proposed fairness constraint. Illustrating a specific scenario will be instrumental in justifying the use and effectiveness of the fairness constraint. The authors state that their formulated fairness constraints are based on the principles of envy-freeness and total welfare. Nevertheless, there may be a potential misapprehension of these principles by the authors as their derived fairness constraints deviate from the standard definitions established from the above notions. Both principles are rooted in the game theory literature and are presented in a context where each game participant possesses a distinct utility function. This utility function is exploited to evaluate the benefits derived from a specific treatment. Envy-freeness and total welfare can be viewed as inherent characteristics of the participants' utility contingent upon their treatments. In essence, the constructs of envy-freeness and total welfare are fundamentally anchored on utility functions. Consequently, it is recommended that the authors shed light on the nature of the utility functions leveraged in their investigation. By doing so, they ensure a more precise understanding of the fairness constraints. In this paper, the authors endeavor to validate Theorems 1 and 2 by invoking Lemma A.1. However, the assertion of Lemma A.1 appears to be flawed, particularly in the analysis of the conditional expectation of $Z_t$ given $\mathcal{F}{t-1}$. This is because $E[Z_t|\mathcal{F}{t-1}]$ is not a random variable and hence cannot equate to $Z_{t-1}$, evident by $E[D_tY_t1(X_t \in \mathcal{S}j)|\mathcal{F}{t-1}] = p_j\hat{e}^*{t,j}E[Y_t|D_t=1,X_t \in \mathcal{S}j]$. However, despite this ostensible flaw in Lemma A.1, it is plausible that the statements of Theorems 1 and 2 might be correct. The validity of these theorems may be independently corroborated by employing the known fact that $Z_t$ possesses a common mean. This subsequently implies that the process of $Z_t$ minus the common mean satisfies the definition of a martingale difference sequence.

Rating

2: Strong Reject: For instance, a paper with major technical flaws, and/or poor evaluation, limited impact, poor reproducibility and mostly unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

1 poor

Limitations

There are no specific limitations or potential negative social impacts that require discussion in relation to the current method.

Reviewer fjNx6/10 · confidence 4/52023-07-07

Summary

This paper is concerned with the problem of experimental design for adaptive experiments under a fairness constraint, i.e., the restriction that some individuals not be disproportionately more likely to receive treatment. The authors approach this problem by modifying the usual Neyman allocation with constraints for envy, welfare, and feasibility of the problem. Update rules are provided for the problem to provide treatment probabilities at each round. The authors also provide theoretical results which bound the ATE and variance, and convergence to the oracle strategy in the limit.

Strengths

* This is a novel task formulation for an important problem. The authors do a nice job of motivating the problem as well, and provide a clear description of the desired characteristics of the solution. * The proposed algorithm is simple and efficient. The authors do a nice job of laying out both problems A and B and clearly describing the update procedure. * Provided theory is nice in being able to reason about the behavior of the proposed procedure. The bound on the ATE and variance is very reasonable, with the proposed procedure taking only a modest penalty in terms of convergence rate. Asymptotic normality and convergence to the optimal strategy in the limit are also nice results.

Weaknesses

While the strategy converges to the optimal strategy in the limit, it's not clear to me how much of an effect these parameters have on the procedure's behavior in finite samples. This is especially important in the design setting, where the main motivation of adaptivity is improving finite sample performance. It would be nice to see some simulated/empirical evidence that examines the robustness of the proposed procedure to misspecification. There is also a natural tension between the welfare and fairness constraints, it would be helpful if the authors provide a discussion of this. It also would have been useful if the authors provided a comparison to a non-adaptive (but balanced) procedure a as an additional baseline (e.g., rerandomizatoin, Gram-Schmidt walk). In this setting, treatment probabilities would remain fixed but variance would still be minimized by accounting for group status with an obvious tradeoff on the welfare constraint.

Questions

* How does this procedure compare with adding regularization to a simple adaptive procedure like UCB? * Little guidance is given for selecting the constraints. In practice how much of an effect do these have on relative performance? Can the authors provide a sense of worst case behavior here?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors overall do a nice job of motivating the problem of fairness in treatment allocation. However, I do think that the introduction of some of these constraints, while well motivated, still leaves a subjective decision in the hands of the experimenter. It seems given this that more attention/discussion should be given to this choice in light of the example used in the paper which has a significant impact on the lives of people who are subjected to the experiment.

Reviewer ZNjb6/10 · confidence 3/52023-07-11

Summary

This paper considers randomized experiments that aim to assess the effectiveness of some treatment or policy. The data and sample efficiency of such randomized experiments has been shown to improve by making them adaptive — i.e. by updating the treatment assignment probabilities on the fly, based on the data collected in the study. This paper argues that this can lead to fairness concerns, however. The goal of this paper is to strike a balance between maximizing information gain through randomized experiments and respecting fairness considerations. The fairness aspects considered in this paper include maximizing overall welfare of the participants (through treatments) and ensuring all participant groups in the experiment receive a “fair” exposure to the treatment. Towards achieving this, the paper proposes a “fair adaptive experimental design strategy” that integrates fairness and welfare considerations along with optimizing information gain. The approach is to build a non-parametric algorithm that calculates the mean and variance of potential outcomes at the group level, each time step, to determine future allocations. The paper presents theoretical results that show that this approach yields strongly consistent estimates. Numerical results show that this approach improves upon the fairness considerations.

Strengths

– I think the key contribution/strength of the paper lies in recognizing and attempting to fill the fairness gap that can arise in adaptive experiments. – The problem formulation and contribution is cleanly spelled out and well-presented. Everything is easy to understand and follow. – The paper presents nice theoretical results backing the soundness of the proposed approach.

Weaknesses

– The experimental section appears a little weak. The DGPs used are completely synthetic and seem to be hand-crafted to demonstrate the utility of this approach. It may be more convincing to have more real/realistic datasets or even synthetic data generated from some real DGP (somewhat similar to the case study, but having one more domain might make it stronger). – The abstract is either too dense or introduces too many technical terms that are discussed only later in the paper, making it very hard to understand. Perhaps cramming less info in the abstract might help. – THe paper doesn’t discuss/ quantify the trade-off involved in achieving fairness: how much information gain is given up as compared to other SOTA adaptive experimentation approaches. Having some insights/ theoretical guarantees here may be very useful.

Questions

– In section 2.1, is $n_t$ the number of *new* participants added to the experiment at time $t$ or simply the total number of participants at time $t$? If it is the former, then why is $D_{it}$ defined only for $i = 1, …, n_t$? If it is the latter, then N= \sum n_t doesn’t seem to make sense (because some participants may get over counted)? Could you please clarify this? – Why is Assumption 1 mild / why is it reasonable or practicable to assume? – In the statement of theorem 1, what is $n_1$ and do you need N tending to infinity? I thought that is already captured by the math? Or am I missing something? – Could you elaborate on the tradeoff involved (data efficiency vs fairness) as mentioned under the last bullet under wekanesses?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

No negative societal impacts discussed (but I don’t see any concerns). Perhaps the technical limitation is that the approach may give up some amount of data efficiency. This could be discussed more in the paper.

Reviewer 62S12023-08-13

Thank you for your detailed rebuttal. After thoroughly reviewing the points made by the authors, my perspective remains largely consistent with my initial review. When comparing this paper to the ones I initially cited, I still find the paper's contribution to be somewhat limited. A more detailed response is provided below. > Contribution compared to existing research I appreciate the authors' emphasis on how this paper distinguishes itself from much of the existing work, specifically by focusing on the design of randomized experiments rather than merely analyzing experimental data to estimate ATE. To my understanding, the central results of this paper hinge on the consistency of the proposed randomized experimental design. The papers I cited earlier also discuss the consistency of arbitrary randomized experimental designs under specific assumptions, which seem to align with the proposed design. The ATE estimation method introduced resembles an AIPW estimator with appropriate hyper-parameters, including weight choices. Therefore, in terms of consistency, the paper seems to be in sync with the analytical scope present in existing research, leading me to view its contribution as relatively modest. While I acknowledge the novelty of the proposed randomized experimental design, its unique aspects appear to be centered around the introduction of specific constraints. Its consistency is recognized, but it largely aligns with existing analyses. The broader utilities of the proposed design seem less explored. For example, the proposed method strives for treatment allocation that adheres to fairness constraints, such as envy-freeness and welfare constraints. However, the resultant optimization problem does not appear to present notable technical challenges. Hence, I perceive this paper as primarily introducing a novel fairness constraint to the randomized experimental design without thoroughly evaluating its benefits beyond consistency. I encourage the authors to delve deeper into the challenges of developing the proposed randomized experimental design. > Assumptions I acknowledge that the proposed randomized experimental design does not depend on the positivity assumption. Yet, there seems to be no fundamental difference between employing a positivity assumption and introducing a feasibility constraint. While the non-requirement of the positivity assumption is highlighted, its significance appears less pronounced. I am surprised that an assumption regarding unmeasured confounders is deemed unnecessary. I recommend that the authors provide a rigorous proof to substantiate this claim. > Envy-freeness From what I understand, envy-freeness does not inherently guarantee a balanced treatment allocation among subgroups. Envy-freeness suggests that an individual's utility exceeds what they might have received from another treatment rather than ensuring a balanced treatment. I'd like to point out that the notion of fairness presented may deviate from a traditional interpretation of envy-freeness.

Authorsrebuttal2023-08-13

We are sorry to hear that you still have concerns about our paper. - You repeatedly mentioned the phrase “consistency of the proposed randomized experimental design,” but it is not a standard terminology in the experimental design literature. As such, we could not understand your comment. - You also suggested that “the resultant optimization problem does not appear to present notable technical challenges.” However, we would like to clarify that we are not trying to deliberately construct a complex optimization algorithm, as this is not the goal of our paper. The goal of our paper is to propose an adaptive experimental design method incorporating fairness and welfare concerns, with solid statistical guarantee (such as asymptotic normality and efficiency). In fact, as randomized field experiments are usually very costly, practitioners tend to avoid procedures that are too complex. For this reason, we believe it is a advantage that our procedure does not involve complicated numerical optimization, a feature that makes our proposal more transparent. - Regarding the “no unmeasured confounders” assumption, we have explained that this assumption holds by construction. In randomized experiments, the treatment assignment is (conditionally) independent of the potential outcomes by construction, and therefore the assumption holds trivially. This is why we do not need to explicitly impose this assumption, and no proof is needed. We note that this is not unique to our paper. For example, Hu and Zhang (2004) is a classical reference. You and another reviewer mentioned the recent work of Simchi-Levi and Wang (2023). None of these two papers had to impose the "no unmeasured confounders" assumption.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC