Hardware Friendly Transformer Optimization with Dynamic Attention Matrix Fusion
Qingyao Yang, Xiaoqin Wang, Yumei Zhou, Qiang Li, Shushan Qiao · 2025
The multi-head self-attention (MHSA) is the core component of the transformer, where dynamic matrix multiplications (DMM), particularly Q×KTand A′×V, pose significant challenges for hardware acceleration. To reduce DMM MACs, this paper proposes a Dynamic Attention Matrix Fusion (DAMF) method, which optimizes DMM from the attention algorithm. For Q × KT, a quadratic form fusion of WQWKweight matrices and an SVD approximation is introduced, transforming DMM into fewer scalar operations and eliminating the linear transformations for QK generation. For A′× V, this paper proposes approximating softmax using a Maclaurin series and power-of-2 as a shift factor, replacing A′× V with a hardware-friendly shift operation. Experimental results show that the proposed DAMF method does not cause significant accuracy loss in the BERT-base. Additionally, compared to MHSA with the same configuration, DAMF reduces parameters by 1.99 times, DMM MACs by 284 times, total MACs by 2.21 times and memory access by 2.71 times.