A Low-Latency and Lightweight FPGA-Based Engine for Softmax and Layer Normalization Acceleration

Seongho Jeong, Minseok Seo, Xuan Truong Nguyen, Hyuk‐Jae Lee · 2023

Vector operations such as GELU, softmax, and layer normalization are essential for transformers, but generally consume long latency on general-purpose CPU and GPU due to their low arithmetic intensities and high nonlinearity. In this study, we propose a low-latency FPGA-based architecture for accelerating the vector operations. Specifically, the proposed design includes processing elements for micro-instructions and a controller that decomposes and executes vector operations into in-order low-level instructions effectively. Experiments show that an FPGA implementation of the proposed design achieves a latency of 1.31us for softmax and 3.76us for layer normalization, while consuming 385 DSPs, 16 BRAMs, 71k FFs, and 57k. LUTs.

Read the paper · More papers on PaperTik