Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature model
Recent evidence has shown the existence of a so-called double-descent and\neven triple-descent behavior for the generalization error of deep-learning\nmodels. This important phenomenon commonly appears in implemented neural\nnetwork architectures, and also seems to emerge in epoch-wise curves during the\ntraining process. A recent line of research has highlighted that random matrix\ntools can be used to obtain precise analytical asymptotics of the\ngeneralization (and training) errors of the random feature model. In this\ncontribution, we analyze the whole temporal behavior of the generalization and\ntraining errors under gradient flow for the random feature model. We show that\nin the asymptotic limit of large system size the full time-evolution path of\nboth errors can be calculated analytically. This allows us to observe how the\ndouble and triple descents develop over time, if and when early stopping is an\noption, and also observe time-wise descent structures. Our techniques are based\non Cauchy complex integral representations of the errors together with recent\nrandom matrix methods based on linear pencils.\n
Paper
References (46)
Scroll for more · 34 remaining