Low DRAM Memory Access and Flexible Dataflow Convolutional Neural Network Accelerator based on RISC-V Custom Instruction

Deep convolutional neural networks has been widely used in several applications. However, the huge computational complexity and data access times hinder its application in edge devices. Previous works target to design a specific fixed dataflow. However, several researches point out that there are no dataflow that can be optimal across all layers or models. In this paper, first we propose an flexible dataflow accelerator which can reconfigure to weight stationary or output stationary dataflow at every layer to increase hardware utilization and data reuse. Besides, we design RISC-V custom instructions to encode the dataflow configurations. Last, we proposed a on-the-fly pooling method to compute the max pooling layer right after the convolutional layer to reduce off-chip memory access. By reconfiguring the dataflow, we improve 3.75x and 1.18x DRAM access amounts in VGG16 compared with [1], [2] respectively. Besides, we maintain a high utilization rate of 99.12%. The proposed accelerator can not only reconfigure the dataflow of each layer but also achieve high throughput, high area efficiency, and high power efficiency. The accelerator implemented in the 40nm process reaching 256 GOPS throughput with 1000 MHz, 136.9 GOPS/mm2 area efficiency with 1.87 mm2 area, 1.014 TOPS/W power efficiency with 252.51 mW power.

Paper

Full text

PDF

Low DRAM Memory Access and Flexible Dataflow Convolutional Neural Network Accelerator based on RISC-V Custom Instruction

Semantic Scholar · Computer Science · 2024

Abstract

Deep convolutional neural networks has been widely used in several applications. However, the huge computational complexity and data access times hinder its application in edge devices. Previous works target to design a specific fixed dataflow. However, several researches point out that there are no dataflow that can be optimal across all layers or models. In this paper, first we propose an flexible dataflow accelerator which can reconfigure to weight stationary or output stationary dataflow at every layer to increase hardware utilization and data reuse. Besides, we design RISC-V custom instructions to encode the dataflow configurations. Last, we proposed a on-the-fly pooling method to compute the max pooling layer right after the convolutional layer to reduce off-chip memory access. By reconfiguring the dataflow, we improve 3.75x and 1.18x DRAM access amounts in VGG16 compared with [1], [2] respectively. Besides, we maintain a high utilization rate of 99.12%. The proposed accelerator can not only reconfigure the dataflow of each layer but also achieve high throughput, high area efficiency, and high power efficiency. The accelerator implemented in the 40nm process reaching 256 GOPS throughput with 1000 MHz, 136.9 GOPS/mm2 area efficiency with 1.87 mm2 area, 1.014 TOPS/W power efficiency with 252.51 mW power.

Similar papers

© 2026 NYSGPT2525 LLC