Understanding Learning Dynamics of Binary Neural Networks via Information Bottleneck

Compact neural networks are essential for affordable and power efficient deep\nlearning solutions. Binary Neural Networks (BNNs) take compactification to the\nextreme by constraining both weights and activations to two levels, $\\{+1,\n-1\\}$. However, training BNNs are not easy due to the discontinuity in\nactivation functions, and the training dynamics of BNNs is not well understood.\nIn this paper, we present an information-theoretic perspective of BNN training.\nWe analyze BNNs through the Information Bottleneck principle and observe that\nthe training dynamics of BNNs is considerably different from that of Deep\nNeural Networks (DNNs). While DNNs have a separate empirical risk minimization\nand representation compression phases, our numerical experiments show that in\nBNNs, both these phases are simultaneous. Since BNNs have a less expressive\ncapacity, they tend to find efficient hidden representations concurrently with\nlabel fitting. Experiments in multiple datasets support these observations, and\nwe see a consistent behavior across different activation functions in BNNs.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC