Several papers argue that wide minima generalize better than narrow minima.\nIn this paper, through detailed experiments that not only corroborate the\ngeneralization properties of wide minima, we also provide empirical evidence\nfor a new hypothesis that the density of wide minima is likely lower than the\ndensity of narrow minima. Further, motivated by this hypothesis, we design a\nnovel explore-exploit learning rate schedule. On a variety of image and natural\nlanguage datasets, compared to their original hand-tuned learning rate\nbaselines, we show that our explore-exploit schedule can result in either up to\n0.84% higher absolute accuracy using the original training budget or up to 57%\nreduced training time while achieving the original reported accuracy. For\nexample, we achieve state-of-the-art (SOTA) accuracy for IWSLT'14 (DE-EN)\ndataset by just modifying the learning rate schedule of a high performing\nmodel.\n
Paper
References (60)
Scroll for more · 38 remaining