On the Effectiveness of Regularization Against Membership Inference Attacks

Deep learning models often raise privacy concerns as they leak information\nabout their training data. This enables an adversary to determine whether a\ndata point was in a model's training set by conducting a membership inference\nattack (MIA). Prior work has conjectured that regularization techniques, which\ncombat overfitting, may also mitigate the leakage. While many regularization\nmechanisms exist, their effectiveness against MIAs has not been studied\nsystematically, and the resulting privacy properties are not well understood.\nWe explore the lower bound for information leakage that practical attacks can\nachieve. First, we evaluate the effectiveness of 8 mechanisms in mitigating two\nrecent MIAs, on three standard image classification tasks. We find that certain\nmechanisms, such as label smoothing, may inadvertently help MIAs. Second, we\ninvestigate the potential of improving the resilience to MIAs by combining\ncomplementary mechanisms. Finally, we quantify the opportunity of future MIAs\nto compromise privacy by designing a white-box `distance-to-confident' (DtC)\nmetric, based on adversarial sample crafting. Our metric reveals that, even\nwhen existing MIAs fail, the training samples may remain distinguishable from\ntest samples. This suggests that regularization mechanisms can provide a false\nsense of privacy, even when they appear effective against existing MIAs.\n

Paper

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC