Thank you for the detailed Ethics Review. In the revision, we will create a dedicated Ethics section to expand on the limitations and ethical concerns already outlined in the paper. We will include a high level outline of why each of these distinct ethical concerns are important, and then delve into details particular to RHEA. Specifically, we will include discussions around each issue brought up in your Recommendation:
*Fairness Constraints.* Thank you for highlighting the importance of this point. As we say in the Broader Impact section of the submission: “In such problems with diverse stakeholders, breaking down costs and benefits by affected populations and allowing users to input explicit constraints to prescriptors can be crucial for generating feasible and equitable models.” For example, the concrete example of a fairness constraint you've suggested could be directly incorporated into the system. RHEA’s multi-objective optimization can use any objective that can be computed based on the system’s behavior, so this fairness objective could be directly used, if impacts on the subgroups can be measured. A human user might also integrate this objective into their calculation of a unified Cost objective, since any deviation from ideal fairness is a societal cost. In deployment, an oversight committee could interrogate any developed metrics before they are used in the optimization process to ensure that they align with declared societal goals as well as possible.
*Governance and Democratic Accountability.* This is a key topic of Project Resilience (reference [1] in the submission), whose goal is to generalize the framework of the Pandemic Response Challenge to SDG goals more broadly. We are currently involved in developing the structure of this platform, in which, for any decision-making project there are three main roles: decision-maker, experts, and the public. The goal is to bring these three roles together under a unified governance structure. At a high-level, the process for any project would be: the Decision-maker defines the problem for which they need help; Experts build models for the problem and makes them (or data to produce Distilled versions) public; The Public comments on the process, including making suggestions what to do in particular contexts, or ways to improve the models (e.g., adding new features, or modifying objectives); Experts incorporate this feedback to update their models; After sufficient discussion, the Decision-maker uses the platform to make decisions, looking at what the public has suggested and what the models suggest, using the Pareto front to make sense of key trade-offs; The Decision-maker communicates about their final decision, i.e., what was considered, why they settled on this set of actions, etc. In this way, key elements of the decision-making process are transparent, and decision-makers can be held accountable for how they integrate this kind of AI system into their decisions. By enabling a public discussion alongside the modeling/optimization process, the system attempts to move AI-assisted decision-making towards participative democracy grounded in science. The closest example of an existing platform with a similar interface is https://www.metaculus.com/home/, but it is for predictions, not prescriptions, and problems are not linked to particular decision-makers.
*Vetting Expert Input.* Some of this vetting happens naturally once the user (i.e. decision-maker) sets specific optimization objectives. For example, if the suggested fairness constraint is used as an objective, intentionally unfair prescriptors can quickly be removed from the population. A unified governance platform like the one outlined above also would enable mechanisms of expert vetting by the public, decision-makers, or other experts.
*Data Privacy and Security.* Since experts submit complete prescriptors, no sensitive data they may have used to build their prescriptors needs to be shared. In the Gather step, each expert team had an independent node to submit their prescriptors. The data for the team was generated by running their prescriptors on their node. The format of the data was then automatically verified, to ensure that it complied with the Defined API. Verified data from all teams was then aggregated for the Distill & Evolve steps. Since the aggregated data all must fit an API that does not allow for extra data to be disclosed, the chance of disclosing sensitive data in the Gather phase is minimized.