Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic Data

Generative models trained with Differential Privacy (DP) can be used to\ngenerate synthetic data while minimizing privacy risks. We analyze the impact\nof DP on these models vis-a-vis underrepresented classes/subgroups of data,\nspecifically, studying: 1) the size of classes/subgroups in the synthetic data,\nand 2) the accuracy of classification tasks run on them. We also evaluate the\neffect of various levels of imbalance and privacy budgets. Our analysis uses\nthree state-of-the-art DP models (PrivBayes, DP-WGAN, and PATE-GAN) and shows\nthat DP yields opposite size distributions in the generated synthetic data. It\naffects the gap between the majority and minority classes/subgroups; in some\ncases by reducing it (a "Robin Hood" effect) and, in others, by increasing it\n(a "Matthew" effect). Either way, this leads to (similar) disparate impacts on\nthe accuracy of classification tasks on the synthetic data, affecting\ndisproportionately more the underrepresented subparts of the data.\nConsequently, when training models on synthetic data, one might incur the risk\nof treating different subpopulations unevenly, leading to unreliable or unfair\nconclusions.\n

Paper

References (66)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC