Taxonomizing local versus global structure in neural network loss landscapes

Viewing neural network models in terms of their loss landscapes has a long\nhistory in the statistical mechanics approach to learning, and in recent years\nit has received attention within machine learning proper. Among other things,\nlocal metrics (such as the smoothness of the loss landscape) have been shown to\ncorrelate with global properties of the model (such as good generalization\nperformance). Here, we perform a detailed empirical analysis of the loss\nlandscape structure of thousands of neural network models, systematically\nvarying learning tasks, model architectures, and/or quantity/quality of data.\nBy considering a range of metrics that attempt to capture different aspects of\nthe loss landscape, we demonstrate that the best test accuracy is obtained\nwhen: the loss landscape is globally well-connected; ensembles of trained\nmodels are more similar to each other; and models converge to locally smooth\nregions. We also show that globally poorly-connected landscapes can arise when\nmodels are small or when they are trained to lower quality data; and that, if\nthe loss landscape is globally poorly-connected, then training to zero loss can\nactually lead to worse test accuracy. Our detailed empirical results shed light\non phases of learning (and consequent double descent behavior), fundamental\nversus incidental determinants of good generalization, the role of load-like\nand temperature-like parameters in the learning process, different influences\non the loss landscape from model and data, and the relationships between local\nand global metrics, all topics of recent interest.\n

Paper

References (97)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC