Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU Networks

A recent line of work has analyzed the theoretical properties of deep neural\nnetworks via the Neural Tangent Kernel (NTK). In particular, the smallest\neigenvalue of the NTK has been related to the memorization capacity, the global\nconvergence of gradient descent algorithms and the generalization of deep nets.\nHowever, existing results either provide bounds in the two-layer setting or\nassume that the spectrum of the NTK matrices is bounded away from 0 for\nmulti-layer networks. In this paper, we provide tight bounds on the smallest\neigenvalue of NTK matrices for deep ReLU nets, both in the limiting case of\ninfinite widths and for finite widths. In the finite-width setting, the network\narchitectures we consider are fairly general: we require the existence of a\nwide layer with roughly order of $N$ neurons, $N$ being the number of data\nsamples; and the scaling of the remaining layer widths is arbitrary (up to\nlogarithmic factors). To obtain our results, we analyze various quantities of\nindependent interest: we give lower bounds on the smallest singular value of\nhidden feature matrices, and upper bounds on the Lipschitz constant of\ninput-output feature maps.\n

Paper

References (60)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC