unzipFPGA: Enhancing FPGA-based CNN Engines with On-the-Fly Weights Generation

Single computation engines have become a popular design choice for FPGA-based\nconvolutional neural networks (CNNs) enabling the deployment of diverse models\nwithout fabric reconfiguration. This flexibility, however, often comes with\nsignificantly reduced performance on memory-bound layers and resource\nunderutilisation due to suboptimal mapping of certain layers on the engine's\nfixed configuration. In this work, we investigate the implications in terms of\nCNN engine design for a class of models that introduce a pre-convolution stage\nto decompress the weights at run time. We refer to these approaches as\non-the-fly. To minimise the negative impact of limited bandwidth on\nmemory-bound layers, we present a novel hardware component that enables the\non-chip on-the-fly generation of weights. We further introduce an input\nselective processing element (PE) design that balances the load between PEs on\nsuboptimally mapped layers. Finally, we present unzipFPGA, a framework to train\non-the-fly models and traverse the design space to select the highest\nperforming CNN engine configuration. Quantitative evaluation shows that\nunzipFPGA yields an average speedup of 2.14x and 71% over optimised status-quo\nand pruned CNN engines under constrained bandwidth and up to 3.69x higher\nperformance density over the state-of-the-art FPGA-based CNN accelerators.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC