An efficient ReRAM-based inference accelerator for convolutional neural networks via activation reuse
In this paper, a novel resistive random access memory (ReRAM) based accelerator is proposed for convolution neural network (CNN) inference accelerations. In ReRAM-based CNN computation, weight parameters can be pre-programmed in ReRAM crossbar arrays, and activations are generated by processing the multiplication-and-accumulation (MAC) operations in the ReRAM crossbar arrays. However, priorworkscannotreuseactivationsincomputation,inwhichtheactivation dominates the data movements and raises significant energy cost. To deal with this dilemma, a tiling-based dataflow is proposed to enable activation reuse among adjacent ReRAM crossbar arrays to reduce the activation movements. We then develop a ReRAM-based CNN accelerator that can well suit the dataflow to reduce the cost of ReRAM access. Evaluation resultsshowthattheproposeddesignachieves1.8 × energysavingand2.8 × bandwidth saving compared with a state-of-the-art PipeLayer accelerator.
Paper
Full text
An efficient ReRAM-based inference accelerator for convolutional neural networks via activation reuse
Semantic Scholar · Computer Science · 2019
Abstract
In this paper, a novel resistive random access memory (ReRAM) based accelerator is proposed for convolution neural network (CNN) inference accelerations. In ReRAM-based CNN computation, weight parameters can be pre-programmed in ReRAM crossbar arrays, and activations are generated by processing the multiplication-and-accumulation (MAC) operations in the ReRAM crossbar arrays. However, priorworkscannotreuseactivationsincomputation,inwhichtheactivation dominates the data movements and raises significant energy cost. To deal with this dilemma, a tiling-based dataflow is proposed to enable activation reuse among adjacent ReRAM crossbar arrays to reduce the activation movements. We then develop a ReRAM-based CNN accelerator that can well suit the dataflow to reduce the cost of ReRAM access. Evaluation resultsshowthattheproposeddesignachieves1.8 × energysavingand2.8 × bandwidth saving compared with a state-of-the-art PipeLayer accelerator.
References (30)
Scroll for more · 18 remaining