Don't Jump Through Hoops and Remove Those Loops: SVRG and Katyusha are Better Without the Outer Loop

The stochastic variance-reduced gradient method (SVRG) and its accelerated\nvariant (Katyusha) have attracted enormous attention in the machine learning\ncommunity in the last few years due to their superior theoretical properties\nand empirical behaviour on training supervised machine learning models via the\nempirical risk minimization paradigm. A key structural element in both of these\nmethods is the inclusion of an outer loop at the beginning of which a full pass\nover the training data is made in order to compute the exact gradient, which is\nthen used to construct a variance-reduced estimator of the gradient. In this\nwork we design {\\em loopless variants} of both of these methods. In particular,\nwe remove the outer loop and replace its function by a coin flip performed in\neach iteration designed to trigger, with a small probability, the computation\nof the gradient. We prove that the new methods enjoy the same superior\ntheoretical convergence properties as the original methods. However, we\ndemonstrate through numerical experiments that our methods have substantially\nsuperior practical behavior.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC