COSA Plus: Enhanced Co-Operative Systolic Arrays for Attention Mechanism in Transformers
Zhican Wang, Gang Wang, Guanghui He · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2024
The attention mechanism is becoming a vital building block across various modern neural networks, e.g., Transformers. However, it encounters low efficiency when deployed on the general-purpose GPU/CPU platform, which motivates the dedicated accelerator design. Existing accelerators are commonly devised by exploring the potential sparsity in attention mechanism using a hardware-software codesign scheme, which suffers from complicated training, fine-tuning processes, and possible accuracy degradation. More importantly, the sparse pattern only focuses on certain datasets with less generality, and the fine-grained sparse pattern could also bring hardware inefficiency. Instead, we try to solve these issues from another perspective: by systematically analysing the inherent dataflow characteristics of the attention mechanism, we propose the co-operative systolic arrays (COSAs) with an optimized dataflow to support the general purpose attention mechanism and pursue higher computational efficiency. COSA system exploits the high parallelism from the inherent model and leverages run-time configurable hybrid dataflows, i.e., weight and output stationary (OS) for a systolic array (SA) to support the varying matrix multiplication in the attention mechanism. Regarding the cascaded matrix multiplications, COSA proposes levels of fusion methodologies to reduce the off-chip access and enhance processing element (PE) utilization, such as directly using the result of OS as the weight of weight stationary SA by deep fusion. Additionally, the COSA system also provides the solution to hide the latency and radically save the buffer size related to the softmax. Experiment results show that, across various benchmarks, COSA can achieve$2.29-2.60\times $throughput improvement over the traditional SA of the same MAC number, with up to 94.7% PE utilization rate and$8.2\times $less off-chip memory access. Compared with the general-purpose platforms,$7.6-12.4\times $energy efficiency over NVIDIA GeForce 3090 GPU and$35.2-80.9\times $energy efficiency over Intel 6226R server CPU.