A Novel Self-Attentive Neural Network Hardware Accelerator for Edge Computing

Fangzhou He, Ke Ding, Xiao He, Mingzhe Chen, Kai Mao, Qian Zhang · 2024

High-precision floating-point self-attentive neural networks have high computational complexity and power consumption, posing significant challenges for deployment on edge devices. To tackle these problems, this paper propose a novel deep neural network hardware accelerator with unique vector tiling systolic array and vector ALU designs, where vector tiling systolic array is used to accelerate matrix multiplications, and vector ALU is used to process non-linear operations. By utilizing a systolic array with a systolic pipeline data path, matrix multiplication is accelerated. The vector tiling mechanism ensures that output results are immediately fed back into the arrays for further accumulation, without additional memory access or data loading delays. This approach significantly reduces arithmetic cycle overhead. Extensive experiments show that our method achieve 70×, 160 ×, 6 × and 9 × increase on peak energy efficiency ratio from Intel i3-2320 CPU, Intel i5-9400 CPU, RTX 2060 GPU and RTX 3060 GPU, respectively. The peak arthmetic of hardware accelerator is 2.41 TOPs in 8-bit fixed-point computation.

Read the paper · More papers on PaperTik