Virtually all state-of-the-art methods for training supervised machine\nlearning models are variants of SGD enhanced with a number of additional\ntricks, such as minibatching, momentum, and adaptive stepsizes. One of the\ntricks that works so well in practice that it is used as default in virtually\nall widely used machine learning software is {\\em random reshuffling (RR)}.\nHowever, the practical benefits of RR have until very recently been eluding\nattempts at being satisfactorily explained using theory. Motivated by recent\ndevelopment due to Mishchenko, Khaled and Richt\\'{a}rik (2020), in this work we\nprovide the first analysis of SVRG under Random Reshuffling (RR-SVRG) for\ngeneral finite-sum problems. First, we show that RR-SVRG converges linearly\nwith the rate $\\mathcal{O}(\\kappa^{3/2})$ in the strongly-convex case, and can\nbe improved further to $\\mathcal{O}(\\kappa)$ in the big data regime (when $n >\n\\mathcal{O}(\\kappa)$), where $\\kappa$ is the condition number. This improves\nupon the previous best rate $\\mathcal{O}(\\kappa^2)$ known for a variance\nreduced RR method in the strongly-convex case due to Ying, Yuan and Sayed\n(2020). Second, we obtain the first sublinear rate for general convex problems.\nThird, we establish similar fast rates for Cyclic-SVRG and Shuffle-Once-SVRG.\nFinally, we develop and analyze a more general variance reduction scheme for\nRR, which allows for less frequent updates of the control variate. We\ncorroborate our theoretical results with suitably chosen experiments on\nsynthetic and real datasets.\n
Paper
References (42)
Scroll for more · 30 remaining