Mutual Information of Neural Network Initialisations: Mean Field\n Approximations

The ability to train randomly initialised deep neural networks is known to\ndepend strongly on the variance of the weight matrices and biases as well as\nthe choice of nonlinear activation. Here we complement the existing geometric\nanalysis of this phenomenon with an information theoretic alternative. Lower\nbounds are derived for the mutual information between an input and hidden layer\noutputs. Using a mean field analysis we are able to provide analytic lower\nbounds as functions of network weight and bias variances as well as the choice\nof nonlinear activation. These results show that initialisations known to be\noptimal from a training point of view are also superior from a mutual\ninformation perspective.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC