EFFICIENT MATRIX MULTIPLICATION ON A PARALLEL PROCESSING DEVICE

Patent №

US 7,792,895

Granted

2010-09-07

Filed 2006

Owner

NVIDIA CORPORATION

AI components

1

hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11454411

The present invention enables efficient matrix multiplication operations on parallel processing devices. One embodiment is a method for mapping CTAs to result matrix tiles for matrix multiplication operations. Another embodiment is a second method for mapping CTAs to result tiles. Yet other embodiments are methods for mapping the individual threads of a CTA to the elements of a tile for result tile computations, source tile copy operations, and source tile copy and transpose operations. The present invention advantageously enables result matrix elements to be computed on a tile-by-tile basis using multiple CTAs executing concurrently on different streaming multiprocessors, enables source tiles to be copied to local memory to reduce the number accesses from the global memory when computing a result tile, and enables coalesced read operations from the global memory as well as write operations to the local memory without bank conflicts.

AI classification

AI hardware1.00
Evolutionary computation0.03
Knowledge representation0.01
Natural language0.00
Machine learning0.00
Speech0.00
Vision0.00
Planning0.00

Ownership

NVIDIA CORPORATION

assignment · 179900800

Assignors

JUFFA, NORBERT, DANILAK, RADOSLAV

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC