Computer vision workloads are increasingly deployed at the edge, driving demand for efficient, reconfigurable accelerators that balance accuracy, performance, resource utilization, and power consumption. While Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) dominate current vision pipelines, the MLP-Mixer is a deviation by utilizing an all-MLP architecture while achieving competitive accuracy. However, the alternating Token- and Channel-mixing structure of MLP-Mixer introduces frequent tensor transpositions, resulting in irregular memory access patterns that pose significant challenges for efficient hardware implementation, particularly on resource-constrained edge platforms.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex