The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks

We study the effects of mild over-parameterization on the optimization\nlandscape of a simple ReLU neural network of the form\n$\\mathbf{x}\\mapsto\\sum_{i=1}^k\\max\\{0,\\mathbf{w}_i^{\\top}\\mathbf{x}\\}$, in a\nwell-studied teacher-student setting where the target values are generated by\nthe same architecture, and when directly optimizing over the population squared\nloss with respect to Gaussian inputs. We prove that while the objective is\nstrongly convex around the global minima when the teacher and student networks\npossess the same number of neurons, it is not even \\emph{locally convex} after\nany amount of over-parameterization. Moreover, related desirable properties\n(e.g., one-point strong convexity and the Polyak-{\\L}ojasiewicz condition) also\ndo not hold even locally. On the other hand, we establish that the objective\nremains one-point strongly convex in \\emph{most} directions (suitably defined),\nand show an optimization guarantee under this property. For the non-global\nminima, we prove that adding even just a single neuron will turn a non-global\nminimum into a saddle point. This holds under some technical conditions which\nwe validate empirically. These results provide a possible explanation for why\nrecovering a global minimum becomes significantly easier when we\nover-parameterize, even if the amount of over-parameterization is very\nmoderate.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC