Deep neural networks can achieve remarkable generalization performances while\ninterpolating the training data perfectly. Rather than the U-curve emblematic\nof the bias-variance trade-off, their test error often follows a "double\ndescent" - a mark of the beneficial role of overparametrization. In this work,\nwe develop a quantitative theory for this phenomenon in the so-called lazy\nlearning regime of neural networks, by considering the problem of learning a\nhigh-dimensional function with random features regression. We obtain a precise\nasymptotic expression for the bias-variance decomposition of the test error,\nand show that the bias displays a phase transition at the interpolation\nthreshold, beyond which it remains constant. We disentangle the variances\nstemming from the sampling of the dataset, from the additive noise corrupting\nthe labels, and from the initialization of the weights. Following up on Geiger\net al. 2019, we first show that the latter two contributions are the crux of\nthe double descent: they lead to the overfitting peak at the interpolation\nthreshold and to the decay of the test error upon overparametrization. We then\nquantify how they are suppressed by ensemble averaging the outputs of K\nindependently initialized estimators. When K is sent to infinity, the test\nerror remains constant beyond the interpolation threshold. We further compare\nthe effects of overparametrizing, ensembling and regularizing. Finally, we\npresent numerical experiments on classic deep learning setups to show that our\nresults hold qualitatively in realistic lazy learning scenarios.\n
Paper
References (60)
Scroll for more · 38 remaining