Low Latency CMOS Hardware Acceleration for Fully Connected Layers in Deep Neural Networks

We present a novel low latency CMOS hardware accelerator for fully connected\n(FC) layers in deep neural networks (DNNs). The FC accelerator, FC-ACCL, is\nbased on 128 8x8 or 16x16 processing elements (PEs) for matrix-vector\nmultiplication, and 128 multiply-accumulate (MAC) units integrated with 128\nHigh Bandwidth Memory (HBM) units for storing the pretrained weights.\nMicro-architectural details for CMOS ASIC implementations are presented and\nsimulated performance is compared to recent hardware accelerators for DNNs for\nAlexNet and VGG 16. When comparing simulated processing latency for a 4096-1000\nFC8 layer, our FC-ACCL is able to achieve 48.4 GOPS (with a 100 MHz clock)\nwhich improves on a recent FC8 layer accelerator quoted at 28.8 GOPS with a 150\nMHz clock. We have achieved this considerable improvement by fully utilizing\nthe HBM units for storing and reading out column-specific FClayer weights in 1\ncycle with a novel colum-row-column schedule, and implementing a maximally\nparallel datapath for processing these weights with the corresponding MAC and\nPE units. When up-scaled to 128 16x16 PEs, for 16x16 tiles of weights, the\ndesign can reduce latency for the large FC6 layer by 60 % in AlexNet and by 3 %\nin VGG16 when compared to an alternative EIE solution which uses compression.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC