How neural networks find generalizable solutions: Self-tuned annealing in deep learning

Despite the tremendous success of Stochastic Gradient Descent (SGD) algorithm\nin deep learning, little is known about how SGD finds generalizable solutions\nin the high-dimensional weight space. By analyzing the learning dynamics and\nloss function landscape, we discover a robust inverse relation between the\nweight variance and the landscape flatness (inverse of curvature) for all\nSGD-based learning algorithms. To explain the inverse variance-flatness\nrelation, we develop a random landscape theory, which shows that the SGD noise\nstrength (effective temperature) depends inversely on the landscape flatness.\nOur study indicates that SGD attains a self-tuned landscape-dependent annealing\nstrategy to find generalizable solutions at the flat minima of the landscape.\nFinally, we demonstrate how these new theoretical insights lead to more\nefficient algorithms, e.g., for avoiding catastrophic forgetting.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC