Qsparse-local-SGD: Distributed SGD with Quantization, Sparsification, and Local Computations

Communication bottleneck has been identified as a significant issue in\ndistributed optimization of large-scale learning models. Recently, several\napproaches to mitigate this problem have been proposed, including different\nforms of gradient compression or computing local models and mixing them\niteratively. In this paper, we propose \\emph{Qsparse-local-SGD} algorithm,\nwhich combines aggressive sparsification with quantization and local\ncomputation along with error compensation, by keeping track of the difference\nbetween the true and compressed gradients. We propose both synchronous and\nasynchronous implementations of \\emph{Qsparse-local-SGD}. We analyze\nconvergence for \\emph{Qsparse-local-SGD} in the \\emph{distributed} setting for\nsmooth non-convex and convex objective functions. We demonstrate that\n\\emph{Qsparse-local-SGD} converges at the same rate as vanilla distributed SGD\nfor many important classes of sparsifiers and quantizers. We use\n\\emph{Qsparse-local-SGD} to train ResNet-50 on ImageNet and show that it\nresults in significant savings over the state-of-the-art, in the number of bits\ntransmitted to reach target accuracy.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC