A 22 nm 10.03-237.99 TOPS/W Time-Digital-Hybrid SRAM Compute-in-Memory AI Accelerator for GNN Edge Device Applications
The utilization of CNN-based accelerators for graph neural network (GNN) applications faces two major challenges: the energy consumption from redundant data movements and the inefficiency in reusing multiply-and-accumulate (MAC) computing. These inadequacies fail to take full advantage of SRAM compute-in-memory (CIM) which is a promising solution to perform MAC operations in artificial intelligence (AI) edge devices. This work overcomes the abovementioned challenges by two chip-level schemes: 1) Neighbor-based Input-Sparsity-Aware Weight-Filtering (NISA-WF) to reduce the data movement, on-chip memory capacity and the number of MAC operations; 2) Repeated-Neighbor-Aware MAC-reuse (RNA-MACR) to further suppress the number of MAC operations and computing latency. We implemented a time-digital-hybrid (TDH) SRAM-CIM design to achieve high accuracy without compromising energy efficiency. A fabricated 22nm AI accelerator chip achieved energy efficiency of 10.03-237.99 TOPS/W with a negligible loss of inference accuracy (< 0.3%) and the TDH SRAM-CIM macro achieved energy efficiency of 20.07-761.57 TOPS/W in 8b mode across various GNN datasets.
Paper
Full text
A 22 nm 10.03-237.99 TOPS/W Time-Digital-Hybrid SRAM Compute-in-Memory AI Accelerator for GNN Edge Device Applications
Semantic Scholar · Computer Science · 2024
Abstract
The utilization of CNN-based accelerators for graph neural network (GNN) applications faces two major challenges: the energy consumption from redundant data movements and the inefficiency in reusing multiply-and-accumulate (MAC) computing. These inadequacies fail to take full advantage of SRAM compute-in-memory (CIM) which is a promising solution to perform MAC operations in artificial intelligence (AI) edge devices. This work overcomes the abovementioned challenges by two chip-level schemes: 1) Neighbor-based Input-Sparsity-Aware Weight-Filtering (NISA-WF) to reduce the data movement, on-chip memory capacity and the number of MAC operations; 2) Repeated-Neighbor-Aware MAC-reuse (RNA-MACR) to further suppress the number of MAC operations and computing latency. We implemented a time-digital-hybrid (TDH) SRAM-CIM design to achieve high accuracy without compromising energy efficiency. A fabricated 22nm AI accelerator chip achieved energy efficiency of 10.03-237.99 TOPS/W with a negligible loss of inference accuracy (< 0.3%) and the TDH SRAM-CIM macro achieved energy efficiency of 20.07-761.57 TOPS/W in 8b mode across various GNN datasets.