In the cloud computing environment, concurrent training of multiple machine learning models will cause serious competition for shared cluster resources and affect the execution efficiency. Aiming at this problem, this paper proposes a cloud computing resource scheduling method for distributed machine learning. Based on historical monitoring data, a model between the number of iterations and model quality improvement is established, the impact of resource allocation on model quality improvement is predicted online, resource optimization scheduling strategies are formulated, and a resource scheduling framework is designed. Experimental results show that the proposed method can quickly adapt to the dynamic changes of tasks and loads and maximize the overall performance of multiple model training jobs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex