Abstract-Present day machine learning is computationallyintensive and processes large amounts of data. It is implementedin a distributed fashion in order to address these scalabilityissues. The work is parallelized across a number of computingnodes. It is usually hard to estimate in advance how manynodes to use for a particular workload. We propose a simpleframework for estimating the scalability of distributed machinelearning algorithms. We measure the scalability by means of thespeedup an algorithm achieves with more nodes. We proposetime complexity models for gradient descent and graphicalmodel inference. We validate the gradient descent model withexperiments on deep learning training and graphical inferenceswith experiments on loopy belief propagation. The proposedframework was used to study the scalability of machine learningalgorithms in Apache Spark
Paper
References (27)
Scroll for more · 15 remaining