BitSys: Bitwise Systolic Array Architecture for Multi-precision Quantized Hardware Accelerators
Yuhao Liu, Salim Ullah, Akash S. Kumar · 2024
Quantized Neural Networks (QNN) have been widely applied in hardware accelerator designs for edge. Because lower precision in quantization leads to higher accuracy loss, the mixed-precision scheme has been explored by using different precision in different layers to trade off resource consumption and inference accuracy. Because regular multiplier designs do not support the reconfiguration for multi-precision, we explored a runtime reconfigurable multi-precision bitwise systolic array design, BitSys, for mixed-precision multiplication in QNN accelerators. The popular design in previous works, such as [1], divides the inputs of multipliers as two parts for four sub-multipliers and preset left shifting, achieving the reconfigurable multiplication by disabling two of the submultipliers. Our design is inspired by the Bitshifter architecture from the works of Liu et al. [2], [1]. We convert the$n\times n$- bit multiplication as$A \times B=\sum_{i=0}^{n-1} \sum_{j=0}^{n-1} 2^{i+j} a_i b_j$. As shown in Figure 1, partial products$P_{i+j}$is the sum of subpartial products$a_{i}b_{j}$with left shifting value,$i+j$. The subpartial product masks select the desired$a_{i}b_{j}$to configure the multi-channel according to the precision. One mask square represents one channel. Therefore, we can implement a bitwise systolic array as Figure 2 (left). The sub-partial product computation and mask are fused in one LUT primitive as one processing element in Figure 2 (right up). The sums of elements, considered the sign-bits, in the diagonal with the same left shifting value shown in Figure 2 (left) are the inputs,$D_{k}$, of the output-generate pipeline of Figure 2 (right down). Systolic array sequentially generates the$D_{k}$, and the final multi-precision output is the sum of all$D_{k}$. We implemented our BitSys multiplier as a systolic array for mixed-precision QNN acceleration. The comparison with the works of Liu et al. [1] is shown in Table I. Our systolic array accelerator consumes more hardware resources than the three single-layer accelerator instances of Liu et al. [1]. However, because of our bitwise processing design, BitSys instance can support 250MHz clock frequency because of the low critical path delay of our multiplier and achieves 188.5-274.7% speed-up in the evaluation of one four-layer 1/2/4/8-bit mixed-precision quantized MLP. Furthermore, our design does not change the input/out width when configuring to different precision, which can be easily integrated into existing accelerator designs.