Provable Memorization via Deep Neural Networks using Sub-linear Parameters

It is known that $O(N)$ parameters are sufficient for neural networks to\nmemorize arbitrary $N$ input-label pairs. By exploiting depth, we show that\n$O(N^{2/3})$ parameters suffice to memorize $N$ pairs, under a mild condition\non the separation of input points. In particular, deeper networks (even with\nwidth $3$) are shown to memorize more pairs than shallow networks, which also\nagrees with the recent line of works on the benefits of depth for function\napproximation. We also provide empirical results that support our theoretical\nfindings.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC