Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss

Neural networks trained to minimize the logistic (a.k.a. cross-entropy) loss\nwith gradient-based methods are observed to perform well in many supervised\nclassification tasks. Towards understanding this phenomenon, we analyze the\ntraining and generalization behavior of infinitely wide two-layer neural\nnetworks with homogeneous activations. We show that the limits of the gradient\nflow on exponentially tailed losses can be fully characterized as a max-margin\nclassifier in a certain non-Hilbertian space of functions. In presence of\nhidden low-dimensional structures, the resulting margin is independent of the\nambiant dimension, which leads to strong generalization bounds. In contrast,\ntraining only the output layer implicitly solves a kernel support vector\nmachine, which a priori does not enjoy such an adaptivity. Our analysis of\ntraining is non-quantitative in terms of running time but we prove\ncomputational guarantees in simplified settings by showing equivalences with\nonline mirror descent. Finally, numerical experiments suggest that our analysis\ndescribes well the practical behavior of two-layer neural networks with ReLU\nactivation and confirm the statistical benefits of this implicit bias.\n

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC