Scheduling deep learning (DL) models to train on powerful clusters with accelerators like GPUs and TPUs, presently falls short, either lacking fine-grained heterogeneity awareness or leaving resources substantially under-utilized. To fill this gap, we propose a novel task-level heterogeneity-aware scheduler for DL clusters, <italic>Hadar</italic>, based on an optimization framework able to boost cluster resource utilization. <italic>Hadar</italic> leverages the performance traits of DL jobs on a heterogeneous DL cluster to make scheduling decisions across both spatial and temporal dimensions. It characterizes the task-level performance heterogeneity for optimization and involves the primal-dual framework employing a dual subroutine, to solve the optimization problem and guide the scheduling design. Our trace-driven simulation with representative DL model training workloads demonstrates that <italic>Hadar</italic> accelerates the total training time duration by 1.20<inline-formula><tex-math notation="LaTeX">$\times$</tex-math><alternatives><mml:math><mml:mo>×</mml:mo></mml:math><graphic position="float" orientation="portrait" xlink:href="chen-ieq1-3672471.gif"/></alternatives></inline-formula> when compared with its state-of-the-art heterogeneity-aware counterpart, Gavel. Further, our <italic>Hadar</italic> scheduler is enhanced to <italic>HadarE</italic> by forking each job into multiple copies to let a job train concurrently on heterogeneous GPUs resided on separate available cluster nodes (i.e., machines or servers) for resource utilization enhancement. <italic>HadarE</italic> is evaluated extensively on physical DL clusters for comparison with <italic>Hadar</italic> and Gavel. With substantial enhancement in cluster resource utilization (by 1.45<inline-formula><tex-math notation="LaTeX">$\times$</tex-math><alternatives><mml:math><mml:mo>×</mml:mo></mml:math><graphic position="float" orientation="portrait" xlink:href="chen-ieq2-3672471.gif"/></alternatives></inline-formula>), <italic>HadarE</italic> exhibits considerable speed-ups in DL model training, reducing the total training time duration by 50% (or 80%) on an Amazon’s AWS (or our lab) cluster, while producing trained DL models with consistently better inference quality than those trained by <italic>Hadar</italic>.
Paper
References (45)
Scroll for more · 33 remaining