Differential privacy provides strong privacy guarantees for machine learning\napplications. Much recent work has been focused on developing differentially\nprivate models, however there has been a gap in other stages of the machine\nlearning pipeline, in particular during the preprocessing phase. Our\ncontributions are twofold: we adapt a privacy violation detection framework\nbased on statistical methods to empirically measure privacy levels of machine\nlearning pipelines, and apply the newly created framework to show that\nresampling techniques used when dealing with imbalanced datasets cause the\nresultant model to leak more privacy. These results highlight the need for\ndeveloping private preprocessing techniques.\n
Paper
References (17)
Scroll for more · 5 remaining