Label noise (stochastic) gradient descent implicitly solves the Lasso for quadratic parametrisation

Understanding the implicit bias of training algorithms is of crucial\nimportance in order to explain the success of overparametrised neural networks.\nIn this paper, we study the role of the label noise in the training dynamics of\na quadratically parametrised model through its continuous time version. We\nexplicitly characterise the solution chosen by the stochastic flow and prove\nthat it implicitly solves a Lasso program. To fully complete our analysis, we\nprovide nonasymptotic convergence guarantees for the dynamics as well as\nconditions for support recovery. We also give experimental results which\nsupport our theoretical claims. Our findings highlight the fact that structured\nnoise can induce better generalisation and help explain the greater\nperformances of stochastic dynamics as observed in practice.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC