A Fully Pipelined Reconfigurable Max–Average Pooling Architecture for FPGA-Based CNN Accelerators

Convolutional neural networks (CNNs) work at a high level as an advanced version of artificial neural networks (ANNs) and are extensively used in image enhancement, object detection, edge AI, and embedded IoT tasks. Pooling operations play a crucial role in CNNs to optimize the spatial dimensions of feature maps. Thus, efficient implementation of pooling operation in hardware is required to improve the overall performance of CNN. Many hardware implementations of pooling in FPGA-based CNN accelerators is executed separately, leading to increased resource usage, operational delay, and limits on the architecture's feasibility. In this paper, we propose a fully pipelined and reconfigurable pooling architecture to support both maximum and average pooling operations within a single unified hardware architecture. The proposed architecture uses the shared comparator and accumulation module, which is supervised by pooling mode control (PM) to switch between max and average pooling modes. Following a systolic array architecture is developed to implement the pipelining and parallelism for streaming data. FPGA implementation of this proposed architecture achieves a <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{1 3. 2} \boldsymbol{\%}$</tex> reduction in hardware resource utilization compared to the existing pooling methods using separate maximum and average pooling operations, and 31.1% higher throughput compared to CMB-Max pooling architecture. The proposed reconfigurable pooling architecture is highly suited for real-time embedded vision and edge AI applications.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC