SQuARM-SGD: Communication-Efficient Momentum SGD for Decentralized Optimization

In this paper, we propose and analyze SQuARM-SGD, a communication-efficient algorithm for decentralized training of large-scale machine learning models over a network. In SQuARM-SGD, each node performs a fixed number of local SGD steps using Nesterov's momentum and then sends sparsified and quantized updates to its neighbors regulated by a locally computable triggering criterion. We provide con…

Paper

Similar papers

© 2026 NYSGPT2525 LLC