Collective Communication Acceleration Architecture for Reduce Operation with Long Operands Based on Efficient Data Access and Reconfigurable On-Chip Buffer
The Reduce operation with long operands is widely used in high-performance computing and AI(Artificial Intelligence) computing. A hardware acceleration architecture for Reduce operation with long operands is designed and implemented in FPGA. Flow control of the accelerator is achieved by using a dedicated hardware trigger mechanism. Based on DMA(Direct Memory Access), an efficient data access method is proposed between host memory and the accelerator. A reconfigurable on-chip buffer structure is designed to achieve flexible and efficient buffering of a large number of operands. Calculations are directly performed on the accelerator with on-chip ALU(Arithmetic Logic Unit) arrays to reduce the communications between the host and NIC(Network Interface Card). Experimental results show that this work has a significant acceleration effect compared to the non-offloading method and the current offloading method in "Tianhe" interconnection.
Paper
Full text
Collective Communication Acceleration Architecture for Reduce Operation with Long Operands Based on Efficient Data Access and Reconfigurable On-Chip Buffer
Semantic Scholar · Computer Science · 2024
Abstract
The Reduce operation with long operands is widely used in high-performance computing and AI(Artificial Intelligence) computing. A hardware acceleration architecture for Reduce operation with long operands is designed and implemented in FPGA. Flow control of the accelerator is achieved by using a dedicated hardware trigger mechanism. Based on DMA(Direct Memory Access), an efficient data access method is proposed between host memory and the accelerator. A reconfigurable on-chip buffer structure is designed to achieve flexible and efficient buffering of a large number of operands. Calculations are directly performed on the accelerator with on-chip ALU(Arithmetic Logic Unit) arrays to reduce the communications between the host and NIC(Network Interface Card). Experimental results show that this work has a significant acceleration effect compared to the non-offloading method and the current offloading method in "Tianhe" interconnection.