Multi-Head Attention Hardware Implementation and Side-Channel Security Analysis for Transformer

Zhiwei Ba, Liji Wu, Jing Hu, Wu Le, Xiangmin Zhang · 2024

This paper proposes an innovative hardware implementation of the complete multi-head attention mechanism, reducing resource consumption and boosting speed while addressing the Transformer model's computational complexity. Experimental results show that the total utilization of logic lookup tables (LUT) is 28,218, accounting for 60% of the total resources, while registers (regs) used are 16,648, accounting for 16% of the total resources. The model with a dimension of 48 was verified on an FPGA development board with the model number xc7k160tfbg676-1. The data of the model was quantized to 16-bit fixed-point, and the final results showed an average error of 0.0069 compared to the software calculations. The execution time was 1.3871 ms, which is 30.26% faster than the 1.9891 ms required by the software. Additionally, by reproducing the typical systolic array structure in matrix operations and implementing side-channel attacks based on its power consumption characteristics, the model structure and parameters were inferred using simple power analysis (SPA) methods.

Read the paper · More papers on PaperTik