Training Linear Neural Networks: Non-Local Convergence and Complexity\n Results

Linear networks provide valuable insights into the workings of neural\nnetworks in general. This paper identifies conditions under which the gradient\nflow provably trains a linear network, in spite of the non-strict saddle points\npresent in the optimization landscape. This paper also provides the\ncomputational complexity of training linear networks with gradient flow. To\nachieve these results, this work develops a machinery to provably identify the\nstable set of gradient flow, which then enables us to improve over the state of\nthe art in the literature of linear networks (Bah et al., 2019;Arora et al.,\n2018a). Crucially, our results appear to be the first to break away from the\nlazy training regime which has dominated the literature of neural networks.\nThis work requires the network to have a layer with one neuron, which subsumes\nthe networks with a scalar output, but extending the results of this\ntheoretical work to all linear networks remains a challenging open problem.\n

Paper

References (80)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC