Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
In the past decade the mathematical theory of machine learning has lagged far\nbehind the triumphs of deep neural networks on practical challenges. However,\nthe gap between theory and practice is gradually starting to close. In this\npaper I will attempt to assemble some pieces of the remarkable and still\nincomplete mathematical mosaic emerging from the efforts to understand the\nfoundations of deep learning. The two key themes will be interpolation, and its\nsibling, over-parameterization. Interpolation corresponds to fitting data, even\nnoisy data, exactly. Over-parameterization enables interpolation and provides\nflexibility to select a right interpolating model.\n As we will see, just as a physical prism separates colors mixed within a ray\nof light, the figurative prism of interpolation helps to disentangle\ngeneralization and optimization properties within the complex picture of modern\nMachine Learning. This article is written with belief and hope that clearer\nunderstanding of these issues brings us a step closer toward a general theory\nof deep learning and machine learning.\n