On the geometry of generalization and memorization in deep neural networks

Understanding how large neural networks avoid memorizing training data is key\nto explaining their high generalization performance. To examine the structure\nof when and where memorization occurs in a deep network, we use a recently\ndeveloped replica-based mean field theoretic geometric analysis method. We find\nthat all layers preferentially learn from examples which share features, and\nlink this behavior to generalization performance. Memorization predominately\noccurs in the deeper layers, due to decreasing object manifolds' radius and\ndimension, whereas early layers are minimally affected. This predicts that\ngeneralization can be restored by reverting the final few layer weights to\nearlier epochs before significant memorization occurred, which is confirmed by\nthe experiments. Additionally, by studying generalization under different model\nsizes, we reveal the connection between the double descent phenomenon and the\nunderlying model geometry. Finally, analytical analysis shows that networks\navoid memorization early in training because close to initialization, the\ngradient contribution from permuted examples are small. These findings provide\nquantitative evidence for the structure of memorization across layers of a deep\nneural network, the drivers for such structure, and its connection to manifold\ngeometric properties.\n

Paper

References (37)

Scroll for more · 25 remaining

Similar papers

© 2026 NYSGPT2525 LLC