Design of an 8-bit floating-point unit for mixed-precision neural network training
X.J. Li, Lu Wang, Guangda Zhang, Xia Zhao, Shiqing Zhang · IET conference proceedings. · 2025
With the rapid growth in the number of participants in neural network models, the demand for storage and computational resources poses a significant challenge for resource-constrained devices. Quantization techniques, as a method of compressing models, have made some progress in the field of 8-bit accuracy, but studies have shown that their applicability is still limited to specific types of networks. In this paper, we first analyze the key factors affecting the training performance of neural networks and propose three fine-grained hybrid accuracy quantization strategies. By finely adjusting the exponent bit-width and mantissa bit-width, these strategies successfully address the problem of neural network training failure under single FP8 precision, improving the training accuracy to near single-precision (FP32) levels. Based on these strategies, this paper designs 8-bit floating-point arithmetic instructions for mixed low-precision training of neural networks on the RISC-V architecture. The instructions support flexible FP8 formats through theewidth3field. We have also designed an 8-bit floating-point unit which primarily supports floating-point operations with specified FP8 precision and introduces a stochastic rounding method. Although the area overhead of our FPU is 35.7% higher than that of traditional FP8 arithmetic unit and the power consumption increases by 8.63%, the fine-grained mixed-precision optimization strategy it adopts significantly improves the accuracy of FP8 training by 81.39%, providing new ideas for the design of low-precision training accelerators and enriching the hardware design of the RISC-V ecosystem.