Bypassing the Ambient Dimension: Private SGD with Gradient Subspace Identification

Differentially private SGD (DP-SGD) is one of the most popular methods for\nsolving differentially private empirical risk minimization (ERM). Due to its\nnoisy perturbation on each gradient update, the error rate of DP-SGD scales\nwith the ambient dimension $p$, the number of parameters in the model. Such\ndependence can be problematic for over-parameterized models where $p \\gg n$,\nthe number of training samples. Existing lower bounds on private ERM show that\nsuch dependence on $p$ is inevitable in the worst case. In this paper, we\ncircumvent the dependence on the ambient dimension by leveraging a\nlow-dimensional structure of gradient space in deep networks -- that is, the\nstochastic gradients for deep nets usually stay in a low dimensional subspace\nin the training process. We propose Projected DP-SGD that performs noise\nreduction by projecting the noisy gradients to a low-dimensional subspace,\nwhich is given by the top gradient eigenspace on a small public dataset. We\nprovide a general sample complexity analysis on the public dataset for the\ngradient subspace identification problem and demonstrate that under certain\nlow-dimensional assumptions the public sample complexity only grows\nlogarithmically in $p$. Finally, we provide a theoretical analysis and\nempirical evaluations to show that our method can substantially improve the\naccuracy of DP-SGD in the high privacy regime (corresponding to low privacy\nloss $\\epsilon$).\n

Paper

References (37)

Scroll for more · 25 remaining

Similar papers

© 2026 NYSGPT2525 LLC