An FPGA-Based Efficient Streaming Vector Processing Engine for Transformer-Based Models

Zicheng He, Tiandong Zhao, Siyuan Miao, Chen Wu, Lei He · 2024

Transformer-based models have obtained extensive success. As their linear operations have been significantly accelerated by a wide range of approaches, nonlinear operations tend to have limited hardware efficiency and become the performance bottleneck. Prior works to accelerate nonlinear operations suffer from either low area efficiency with poor instruction set architecture (ISA) or limited throughput due to the loop-carried dependency in reduce operations. In this paper, we propose a vector processing engine (VPE) with new streaming ISA to achieve flexible streaming execution for nonlinear operations and obtain better hardware efficiency. Moreover, we relax the loop-carried dependency in reduce operations with look-ahead dataflow optimization and improve the throughput of nonlinear operations. Experimental results on Xilinx U200 FPGA show that VPE can outperform other FPGA vector processing units on throughput by 1.14 x -3 x for softmax, layer normalization, GELU across different vector sizes.

Read the paper · More papers on PaperTik