PREMA: A Predictive Multi-task Scheduling Algorithm For Preemptible Neural Processing Units

To amortize cost, cloud vendors providing DNN acceleration as a service to\nend-users employ consolidation and virtualization to share the underlying\nresources among multiple DNN service requests. This paper makes a case for a\n"preemptible" neural processing unit (NPU) and a "predictive" multi-task\nscheduler to meet the latency demands of high-priority inference while\nmaintaining high throughput. We evaluate both the mechanisms that enable NPUs\nto be preemptible and the policies that utilize them to meet scheduling\nobjectives. We show that preemptive NPU multi-tasking can achieve an average\n7.8x, 1.4x, and 4.8x improvement in latency, throughput, and SLA satisfaction,\nrespectively.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC