Network Contention-Aware Cluster Scheduling with Reinforcement Learning

Network contention can significantly degrade training throughput of deep learning jobs in GPU clusters. In this paper, we present a new approach to mitigate network contention by formulating GPU cluster scheduling as a reinforcement learning problem. We show that compared to widely used scheduling policies, our approach reduces average job completion time by up to 18.2% and effectively cuts the tail job completion time by up to 20.7%, while allowing a preferable trade-off with resource utilization. More details can be found in our full paper [1] and we open our work at https://github.com/gajagajago/deepshare.

Paper

Similar papers

© 2026 NYSGPT2525 LLC