A novel model quantization scheme for efficient deep neural networks inference and training
Sung-En Chang · 2024
DNNs have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference and training on edge-computing platforms, e.g., ASICs, FPGAs, and embedded systems, are extensively investigated. Due to the considerable model size and computation amount, model compression is a critical step in deploying DNN models on edge devices. This work focuses on model quantization, a hardware-friendly model compression approach that is helpful for efficient inference and efficient training. For efficient inference, we propose a hardware-friendly quantization scheme named sum-of-power-of-2 (SP2), in which the multiplication arithmetic can be replaced with a logic shifter and adder, thereby enabling highly efficient implementations with the FPGA LUT resources. In contrast, the existing fixed-point quantization only can be implemented efficiently by DSP, which is a relatively scarce computing resource of the FPGA. Further, we propose RMSMP, with a Row-wise Mixed-Scheme and Multi-Precision approach. Specifically, we assign mixed quantization schemes and multiple precisions within layers - among rows of the DNN weight matrix, for simplified operations in hardware inference, while preserving accuracy. Further, we found that the quantization error does not necessarily exhibit the layer-wise sensitivity, and actually can be mitigated as long as a certain portion of the weights in every layer are in higher precisions. This observation enables layer-wise uniformity in the hardware implementation towards guaranteed inference acceleration, while still enjoying row-wise flexibility of mixed schemes and multiple precisions to boost accuracy. The candidates of schemes and precisions are derived practically and effectively with a highly hardware-informative strategy to reduce the problem search space. For efficient training, stochastic rounding is crucial in the low-bit (e.g., 8-bit) training of deep neural networks (DNNs) to achieve high accuracy. One of the drawbacks of prior studies is that they require a large number of high-precision stochastic rounding units (SRUs) to guarantee l ow-bit DNN accuracy, which involves considerable hardware overhead. Thus, we use extremely low-bit SRUs (ESRUs) to save a large number of hardware resources during low-bit DNN training. However, a naively designed ESRU introduces a biased distribution of random numbers, causing accuracy degradation. To address this issue, we further propose an ESRU design with a plateau-shape distribution. The plateau-shape distribution in our ESRU design is implemented with the combination of an LFSR (linear-feedback shift register) and an inverted LFSR, which turns an inherent LFSR drawback into an advantage in our efficient ESRU design.--Author's abstract