MPICC: Multiple-Precision Inter-Combined MAC Unit with Stochastic Rounding for Ultra-Low-Precision Training

Leran Huang, Yongpan Liu, Xinyuan Lin, Chenhan Wei, Wenyu Sun, Zengwei Wang, Boran Cao, Chi Zhang, Xiaoxia Fu, Wentao Zhao, Sheng Zhang · 2025

Recent studies have proved the feasibility of ultra-low-precision (≤ 8-bit) training. However, most of the existing operational circuits support only a few higher precisions (FP16, FP32, etc.) for multiplication, and the bit width for accumulation cannot be reduced, which results in low area and power efficiencies. In this paper, we propose MPICC, a multiple-precision Multiply-Accumulate (MAC) unit designed for ultra-low-precision training. It supports inter-combined computations across 18 different precision combinations, including LOG4 (radix-4 FP4), FP4, FP6, FP8, and INT4. It also reduces the accumulation precision from FP16 to FP12 through an optimized Stochastic Rounding (SR) strategy to further save logic resources. Moreover, a low-cost emulating controller, which time-division multiplexed the low-precision MAC unit, is also designed to accomplish high-precision computations for critical DNN layers. Compared with the existing multiple-precision computing units, the area/energy efficiencies of this design are improved by 1.17×/1.19× at FP8, and 4.69×/3.64× at FP4, respectively. The SR strategy further reduces the area/power consumption by 15.6%/14.9% of the floating-point accumulator.

Read the paper · More papers on PaperTik