Disaggregated evaluations of AI systems, in which system performance is\nassessed and reported separately for different groups of people, are\nconceptually simple. However, their design involves a variety of choices. Some\nof these choices influence the results that will be obtained, and thus the\nconclusions that can be drawn; others influence the impacts -- both beneficial\nand harmful -- that a disaggregated evaluation will have on people, including\nthe people whose data is used to conduct the evaluation. We argue that a deeper\nunderstanding of these choices will enable researchers and practitioners to\ndesign careful and conclusive disaggregated evaluations. We also argue that\nbetter documentation of these choices, along with the underlying considerations\nand tradeoffs that have been made, will help others when interpreting an\nevaluation's results and conclusions.\n
Paper
References (75)
Scroll for more · 38 remaining