PIPELINED APPROACH TO FUSED KERNELS FOR OPTIMIZATION OF MACHINE LEARNING WORKLOADS ON GRAPHICAL PROCESSING UNITS
Patent №
US 9,972,063
Granted
2018-05-15
Filed 2015
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
2
ml · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
14813522
A method for optimization of machine learning (ML) workloads on a graphics processor unit (GPU). The method includes identifying a computation having a generic pattern commonly observed in ML processes. An optimized fused GPU kernel is employed to exploit temporal locality for inherent data-flow dependencies in the identified computation. Hierarchical aggregation spanning a memory hierarchy of the GPU for processing for the identified computation is performed. GPU kernel launch parameters are estimated following an analytical model that maximizes thread occupancy and minimizes atomic writes to GPU global memory.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 365050150
Assignors
ASHARI, ARASH, BOEHM, MATTHIAS, CAMPBELL, KEITH W., EVFIMIEVSKI, ALEXANDRE V., KEENLEYSIDE, JOHN D., REINWALD, BERTHOLD, TATIKONDA, SHIRISH
On an employer assignment, the assignors are typically the inventors.