MPN-S: An Algorithm-Hardware Co-Designed Architecture for High-Efficiency Softmax Computation

Yuchen Ma, Defa Wu, Yucheng Du, Xing Wang, Xin Si, Chao Chen · IEEE Transactions on Circuits & Systems II Express Briefs · 2025

As a fundamental operator in Transformer, softmax presents critical hardware challenges due to its computational complexity. The wide use of exponential and division operations in softmax not only occupies a large amount of computing resources, but also causes a high execution delay in hardware implementation. This work presents a high-precision algorithm and hardware co-designed softmax architecture including Mapping optimization, dynamic Pruning, and dynamic Newton iteration method with pipeline Scheduling(MPN-S). The proposed MPN-S simplifies the softmax calculation through the minimum error integer mapping method, dynamic pruning and compressed lookup tables. Combined with dynamic Newton iteration and pipeline design, MPN-S compresses the softmax processing time into the time of finding the maximum value once. The implementation results show that, under 28 nm CMOS technology at the frequency of 0.5 GHz, MPN-S can achieve the efficiency of 3827.18 Gps/(mm2∙ mW) with the area of 2292 μm2 excluding buffers. Neural network experiments show that after replacing the softmax function with MPN-S, the accuracy rate on the MobileVIT classification task reaches 92.9%, and the mAP50val on Yolo V10 reaches 0.959.

Read the paper · More papers on PaperTik