Understanding the implicit bias of training algorithms is of crucial\nimportance in order to explain the success of overparametrised neural networks.\nIn this paper, we study the dynamics of stochastic gradient descent over\ndiagonal linear networks through its continuous time version, namely stochastic\ngradient flow. We explicitly characterise the solution chosen by the stochastic\nflow and prove that it always enjoys better generalisation properties than that\nof gradient flow. Quite surprisingly, we show that the convergence speed of the\ntraining loss controls the magnitude of the biasing effect: the slower the\nconvergence, the better the bias. To fully complete our analysis, we provide\nconvergence guarantees for the dynamics. We also give experimental results\nwhich support our theoretical claims. Our findings highlight the fact that\nstructured noise can induce better generalisation and they help explain the\ngreater performances observed in practice of stochastic gradient descent over\ngradient descent.\n