Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training

We study the effect of mini-batching on the loss landscape of deep neural\nnetworks using spiked, field-dependent random matrix theory. We demonstrate\nthat the magnitude of the extremal values of the batch Hessian are larger than\nthose of the empirical Hessian. We also derive similar results for the\nGeneralised Gauss-Newton matrix approximation of the Hessian. As a consequence\nof our theorems we derive an analytical expressions for the maximal learning\nrates as a function of batch size, informing practical training regimens for\nboth stochastic gradient descent (linear scaling) and adaptive algorithms, such\nas Adam (square root scaling), for smooth, non-convex deep neural networks.\nWhilst the linear scaling for stochastic gradient descent has been derived\nunder more restrictive conditions, which we generalise, the square root scaling\nrule for adaptive optimisers is, to our knowledge, completely novel. %For\nstochastic second-order methods and adaptive methods, we derive that the\nminimal damping coefficient is proportional to the ratio of the learning rate\nto batch size. We validate our claims on the VGG/WideResNet architectures on\nthe CIFAR-$100$ and ImageNet datasets. Based on our investigations of the\nsub-sampled Hessian we develop a stochastic Lanczos quadrature based on the fly\nlearning rate and momentum learner, which avoids the need for expensive\nmultiple evaluations for these key hyper-parameters and shows good preliminary\nresults on the Pre-Residual Architecure for CIFAR-$100$.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC