PIPELINED APPROACH TO FUSED KERNELS FOR OPTIMIZATION OF MACHINE LEARNING WORKLOADS ON GRAPHICAL PROCESSING UNITS

Patent №

US 9,972,063

Granted

2018-05-15

Filed 2015

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

2

ml · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14813522

A method for optimization of machine learning (ML) workloads on a graphics processor unit (GPU). The method includes identifying a computation having a generic pattern commonly observed in ML processes. An optimized fused GPU kernel is employed to exploit temporal locality for inherent data-flow dependencies in the identified computation. Hierarchical aggregation spanning a memory hierarchy of the GPU for processing for the identified computation is performed. GPU kernel launch parameters are estimated following an analytical model that maximizes thread occupancy and minimizes atomic writes to GPU global memory.

AI classification

Machine learning1.00
AI hardware1.00
Planning0.20
Vision0.02
Knowledge representation0.01
Natural language0.00
Speech0.00
Evolutionary computation0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 365050150

Assignors

ASHARI, ARASH, BOEHM, MATTHIAS, CAMPBELL, KEITH W., EVFIMIEVSKI, ALEXANDRE V., KEENLEYSIDE, JOHN D., REINWALD, BERTHOLD, TATIKONDA, SHIRISH

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC