Optimizing Temporal Convolutional Network inference on FPGA-based\n accelerators

Convolutional Neural Networks are extensively used in a wide range of\napplications, commonly including computer vision tasks like image and video\nclassification, recognition, and segmentation. Recent research results\ndemonstrate that multilayer(deep) networks involving mono-dimensional\nconvolutions and dilation can be effectively used in time series and sequences\nclassification and segmentation, as well as in tasks involving sequence\nmodelling. These structures, commonly referred to as Temporal Convolutional\nNetworks (TCNs), have been demonstrated to consistently outperform Recurrent\nNeural Networks in terms of accuracy and training time [1]. While FPGA-based\ninference accelerators for classic CNNs are widespread, literature is lacking\nin a quantitative evaluation of their usability on inference for TCN models. In\nthis paper we present such an evaluation, considering a CNN accelerator with\nspecific features supporting TCN kernels as a reference and a set of\nstate-of-the-art TCNs as a benchmark. Experimental results show that, during\nTCN execution, operational intensity can be critical for the overall\nperformance. We propose a convolution scheduling based on batch processing that\ncan boost efficiency up to 96% of theoretical peak performance. Overall we can\nachieve up to 111,8 GOPS/s and power efficiency of 33,9 GOPS/s/W on an\nUltrascale+ ZU3EG (up to 10x speedup and 3x power efficiency improvement with\nrespect to pure software implementation).\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC