On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning

Machine unlearning, i.e. having a model forget about some of its training\ndata, has become increasingly more important as privacy legislation promotes\nvariants of the right-to-be-forgotten. In the context of deep learning,\napproaches for machine unlearning are broadly categorized into two classes:\nexact unlearning methods, where an entity has formally removed the data point's\nimpact on the model by retraining the model from scratch, and approximate\nunlearning, where an entity approximates the model parameters one would obtain\nby exact unlearning to save on compute costs. In this paper, we first show that\nthe definition that underlies approximate unlearning, which seeks to prove the\napproximately unlearned model is close to an exactly retrained model, is\nincorrect because one can obtain the same model using different datasets. Thus\none could unlearn without modifying the model at all. We then turn to exact\nunlearning approaches and ask how to verify their claims of unlearning. Our\nresults show that even for a given training trajectory one cannot formally\nprove the absence of certain data points used during training. We thus conclude\nthat unlearning is only well-defined at the algorithmic level, where an\nentity's only possible auditable claim to unlearning is that they used a\nparticular algorithm designed to allow for external scrutiny during an audit.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC