Deep learning (DL) constitutes a significant workload within public or private cloud platforms. Previous studies have demonstrated that communication overhead of interconnected GPU clusters could be a significant component of the overall DL (training) time, and thus minimizing the communication overhead with intelligent scheduling of jobs on physically close GPUs should significantly reduce DL time. However, different Deep Neural Network models display varied communication overhead even with similar levels of consolidation. We propose a novel GPU-cluster scheduler, Dally, for Distributed DL (DDL) workloads, that enables proximity-based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication overhead. Dally consists of three major components: (i) a mechanism based on the classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an “autotuner” to optimize delay timers for effective delay scheduling. Furthermore, to enable a cost-effective methodology for largescale experiments, we develop a data-driven DDL cluster simulation platform ArtISt-sim. Dally can provide an improvement in makespan of up to $69 \%(68 \%$ mean) in all training jobs, compared to the prevailing consolidation-based scheduling methods. The resultant increase in system throughput enables the reduction of the average job completion time by up to 36% (26 % mean) and decreases the average communication overhead by up to 83 % (66 % mean), under congested network conditions, compared to the current state of the art.
Paper
Full text
Dally: A Network-Placement Sensitive Cluster Scheduler for Deep Learning
Semantic Scholar · Computer Science · 2025
Abstract
Deep learning (DL) constitutes a significant workload within public or private cloud platforms. Previous studies have demonstrated that communication overhead of interconnected GPU clusters could be a significant component of the overall DL (training) time, and thus minimizing the communication overhead with intelligent scheduling of jobs on physically close GPUs should significantly reduce DL time. However, different Deep Neural Network models display varied communication overhead even with similar levels of consolidation. We propose a novel GPU-cluster scheduler, Dally, for Distributed DL (DDL) workloads, that enables proximity-based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication overhead. Dally consists of three major components: (i) a mechanism based on the classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an “autotuner” to optimize delay timers for effective delay scheduling. Furthermore, to enable a cost-effective methodology for largescale experiments, we develop a data-driven DDL cluster simulation platform ArtISt-sim. Dally can provide an improvement in makespan of up to $69 %(68 %$ mean) in all training jobs, compared to the prevailing consolidation-based scheduling methods. The resultant increase in system throughput enables the reduction of the average job completion time by up to 36% (26 % mean) and decreases the average communication overhead by up to 83 % (66 % mean), under congested network conditions, compared to the current state of the art.