FeCaffe: FPGA-enabled Caffe with OpenCL for Deep Learning Training and Inference on Intel Stratix 10
Deep learning has becoming increasingly more popular in recent years, and there are many popular frameworks in the market accordingly, such as Caffe, TensorFlow and Pytorch. All these frameworks natively support CPUs and GPGPUs. However, FPGAs still cannot provide a comprehensive support by these frameworks for deep learning development, especially for the training phase. In this paper, we firstly propose the FeCaffe, i.e. FPGA-enabled Caffe, a hierarchical software and hardware design methodology based on the Caffe, to enable FPGA to support CNN training features. Moreover, we provide some benchmarks of popular CNN networks with FeCaffe, and further analysis in details accordingly. Finally, some optimization directions including FPGA kernel design, system pipeline, network architecture, user case application and heterogeneous platform levels, have been proposed gradually. The result demonstrates the proposed FeCaffe can support almost full features for training and inference respectively with high degree of design flexibility, expansibility and reusability for deep learning development. Compared to prior studies, our architecture can support more network and training settings and current configuration can achieve 6.4x and 8.4x average execution time improvement for forward and backward respectively for LeNet.
Paper
References (26)
Scroll for more · 14 remaining