The high computational demand of deep neural networks puts forward the need for deep model quantization, to deploy in resource-constrained devices. This paper proposes a Spatial-Depth layer wise Mixed-Precision Quantization scheme that analyzes the structural type and sensitivity of each layer to determine its importance to model performance, thus guiding the selection of appropriate bit-width. The proposed method analyzed the sensitivity of layers by isolating the quantization effect on two types of layers, i.e., spatial feature extractors and depth feature extractors, in Pseudo 3D Convolutional Neural Networks (P3DCNNs). By isolating layer-type quantization, the corresponding gradient behavior and performance drop are leveraged to design a mixed precision quantization strategy. This approach gives a per-layer sensitivity score that captures both local behavior through gradients and global significance based on layer-type performance impact. Experiments are conducted on BraTS-2020 dataset, and the results demonstrate the effectiveness of SD-LAMQuant for mixed precision bit-width assignment.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex