Activation function design for deep networks: linearity and effective initialisation

The activation function deployed in a deep neural network has great influence\non the performance of the network at initialisation, which in turn has\nimplications for training. In this paper we study how to avoid two problems at\ninitialisation identified in prior works: rapid convergence of pairwise input\ncorrelations, and vanishing and exploding gradients. We prove that both these\nproblems can be avoided by choosing an activation function possessing a\nsufficiently large linear region around the origin, relative to the bias\nvariance $\\sigma_b^2$ of the network's random initialisation. We demonstrate\nempirically that using such activation functions leads to tangible benefits in\npractice, both in terms test and training accuracy as well as training time.\nFurthermore, we observe that the shape of the nonlinear activation outside the\nlinear region appears to have a relatively limited impact on training. Finally,\nour results also allow us to train networks in a new hyperparameter regime,\nwith a much larger bias variance than has previously been possible.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC