FAS-Trans: Fully Exploiting FFN and Attention Sparsity for Transformer on FPGA

Hongji Wang, Kun Wang, Yifan Zhang, Jun Wei Yu · 2024

Transformers have increasingly become the backbone of modern AI, excelling in tasks across natural language processing (NLP) and computer vision (CV). However, deploying them on resource-constrained platforms is challenging due to their high computational and energy demands. Previous efforts primarily focused on reducing the computational load of the self-attention module in Transformers, often neglecting optimization for other parts, like feed-forward network (FFN) modules. To address this gap, we propose FAS-Trans, an innovative algorithm-architecture co-design accelerator that efficiently optimizes both self-attention and FFN modules. FAS-Trans incorporates an innovative approximate prediction mechanism utilizing shifted-adders, which pre-estimates matrix sparsity to further reduce computational loads. This mechanism also enables reusing approximate prediction values in subsequent exact computations. Moreover, our approach includes cross-stage sparsity prediction for self-attention module to minimize computations involved in both QKV generation and attention computing. In the FFN module, we predict and exploit sparsity of FC1 block and employ low-precision multipliers for values close to zero, significantly cutting down FFN computational overhead. Our dedicated hardware architecture can effectively handle the irregularities of sparsity and multi-precision, ensuring high hardware resource utilization. Comprehensive evaluations across multiple benchmarks demonstrate that FAS-Trans reduces normalized computational overhead by 41.9% on average with 1% loss in accuracy. FAS-Trans can achieve 2.39-27.98× speedup and 18.9--72.6× energy efficiency improvement compared with CPU and GPU acceleration. Furthermore, we observe that FAS-Trans achieves 2.15--3.80× speedup, 1.87--19.89× improvement in throughput and 2.4--5.6× improvement in energy efficiency compared with other FPGA-based Transformer accelerators.

Read the paper · More papers on PaperTik