Hardware Friendly Transformer Optimization with Dynamic Attention Matrix Fusion

Qingyao Yang, Xiaoqin Wang, Yumei Zhou, Qiang Li, Shushan Qiao · 2025

The multi-head self-attention (MHSA) is the core component of the transformer, where dynamic matrix multiplications (DMM), particularly Q×KTand A′×V, pose significant challenges for hardware acceleration. To reduce DMM MACs, this paper proposes a Dynamic Attention Matrix Fusion (DAMF) method, which optimizes DMM from the attention algorithm. For Q × KT, a quadratic form fusion of WQWKweight matrices and an SVD approximation is introduced, transforming DMM into fewer scalar operations and eliminating the linear transformations for QK generation. For A′× V, this paper proposes approximating softmax using a Maclaurin series and power-of-2 as a shift factor, replacing A′× V with a hardware-friendly shift operation. Experimental results show that the proposed DAMF method does not cause significant accuracy loss in the BERT-base. Additionally, compared to MHSA with the same configuration, DAMF reduces parameters by 1.99 times, DMM MACs by 284 times, total MACs by 2.21 times and memory access by 2.71 times.

Read the paper · More papers on PaperTik