ReLU neural networks define piecewise linear functions of their inputs. However, initializing and training a neural network is very different from fitting a linear spline. In this paper, we expand empirically upon previous theoretical work to demonstrate features of trained neural networks. Standard network initialization and training produce networks vastly simpler than a naive parameter count would suggest and can impart odd features to the trained network. However, we also show the forced simplicity is beneficial and, indeed, critical for the wide success of these networks.
Paper
References (8)
01The Upper Bound on Knots in Neural NetworksKevin K. Chen2016 · arXiv.org · 15 citations In Library
07URL: https://github.com/fchollet/keras2015 · François Chollet. Keras. Github
08Xavier Glorot and Yoshua Bengio Understanding the difficulty of training deep feedforward neural networks, International Conference on Artificial Intelligence and Statistics2010 · Xavier Glorot and Yoshua Bengio Understanding the difficulty of training deep feedforward neural networks, International Conference on Artificial Intelligence and Statistics