An Empirical Comparison of Bias Reduction Methods on Real-World Problems in High-Stakes Policy Settings

Applications of machine learning (ML) to high-stakes policy settings -- such\nas education, criminal justice, healthcare, and social service delivery -- have\ngrown rapidly in recent years, sparking important conversations about how to\nensure fair outcomes from these systems. The machine learning research\ncommunity has responded to this challenge with a wide array of proposed\nfairness-enhancing strategies for ML models, but despite the large number of\nmethods that have been developed, little empirical work exists evaluating these\nmethods in real-world settings. Here, we seek to fill this research gap by\ninvestigating the performance of several methods that operate at different\npoints in the ML pipeline across four real-world public policy and social good\nproblems. Across these problems, we find a wide degree of variability and\ninconsistency in the ability of many of these methods to improve model\nfairness, but post-processing by choosing group-specific score thresholds\nconsistently removes disparities, with important implications for both the ML\nresearch community and practitioners deploying machine learning to inform\nconsequential policy decisions.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC