Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates

It has been experimentally observed that the efficiency of distributed\ntraining with stochastic gradient (SGD) depends decisively on the batch size\nand -- in asynchronous implementations -- on the gradient staleness.\nEspecially, it has been observed that the speedup saturates beyond a certain\nbatch size and/or when the delays grow too large. We identify a data-dependent\nparameter that explains the speedup saturation in both these settings. Our\ncomprehensive theoretical analysis, for strongly convex, convex and non-convex\nsettings, unifies and generalized prior work directions that often focused on\nonly one of these two aspects. In particular, our approach allows us to derive\nimproved speedup results under frequently considered sparsity assumptions. Our\ninsights give rise to theoretically based guidelines on how the learning rates\ncan be adjusted in practice. We show that our results are tight and illustrate\nkey findings in numerical experiments.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC