Matrix multiplication is the dominant computation during Machine Learning (ML) inference. To efficiently perform such multiplication operations, Compute-in-memory (CiM) paradigms have emerged as a highly energy efficient solution. However, integrating compute in memory poses key questions, such as 1) <italic>What type of CiM to use:</italic> Given a multitude of CiM design characteristics, determining their suitability from architecture perspective is needed. 2) <italic>When to use CiM:</italic> ML inference includes workloads with a variety of memory and compute requirements, making it difficult to identify when CiM is more beneficial than standard processing cores. 3) <italic>Where to integrate CiM:</italic> Each memory level has different bandwidth and capacity, creating different data reuse opportunities for CiM integration. To answer such questions regarding on-chip CiM integration for accelerating ML workloads, we use an analytical architecture-evaluation methodology with tailored mapping algorithm. The mapping algorithm aims to achieve highest weight reuse and reduced data movements for a given CiM prototype and workload. Our analysis considers the integration of CiM prototypes into the cache levels of a tensor-core-like architecture, and shows that CiM integrated memory improves energy efficiency by up to <inline-formula><tex-math notation="LaTeX">$3.4 \times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>3</mml:mn><mml:mo>.</mml:mo><mml:mn>4</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="sharma-ieq1-3574508.gif"/></alternatives></inline-formula> and throughput by up to <inline-formula><tex-math notation="LaTeX">$15.6 \times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>15</mml:mn><mml:mo>.</mml:mo><mml:mn>6</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="sharma-ieq2-3574508.gif"/></alternatives></inline-formula> compared to established baseline with INT-8 precision. We believe the proposed work provides insights into <italic>what</italic> type of CiM to use, and <italic>when</italic> and <italic>where</italic> to optimally integrate it in the cache hierarchy for efficient matrix multiplication.
Paper
References (48)
Scroll for more · 36 remaining