Usable Information and Evolution of Optimal Representations During Training

We introduce a notion of usable information contained in the representation learned by a deep network, and use it to study how optimal representations for the task emerge during training, and how they adapt to different tasks. We use this to characterize the transient dynamics of deep neural networks on perceptual decision-making tasks inspired by neuroscience literature, as well as on standard image classification tasks. We show that both the random initialization and the implicit regularization from Stochastic Gradient Descent play an important role in learning minimal sufficient representations for the task. In addition, we evaluate how perturbing the initial part of training impacts the learning dynamics and resulting representations.

Paper

Similar papers

© 2026 NYSGPT2525 LLC