An Efficient FPGA-Based Hardware Accelerator of Fully Quantized Mamba-2

Kailing Zhou, Han Jiao, Wenjin Huang, Yihua Huang · 2025

The Mamba-2 model introduces a State Space Duality (SSD) mechanism, based on original State Space Models (SSMs), that accelerates training and improves accuracy. However, efficient hardware acceleration for Mamba-2 faces challenges. Numerous element-wise operations fail to fully utilize GPU tensor cores, diminishing inference efficiency. Furthermore, research on full-quantization strategies for Mamba-2 is lacking. To address this, we propose a hybrid-precision full-quantization strategy, Hfqmamba2, balancing performance and hardware resource usage. Applying this strategy to Mamba-2 models of various sizes shows that accuracy loss remains within an acceptable range. We also propose an efficient FPGA-based hardware accelerator for Mamba-2. Given the distinct data flow characteristics of the two RMSNorm (Root Mean Square Normalization) layers in the hardware implementation, we introduce a reconfigurable hardware architecture based on a segmented quantization strategy, improving efficiency and flexibility by using a segmented lookup table to approximate the inverse square root operation. For the selective SSM layer operations, we design an intra-layer computation pipeline to enhance processing efficiency. Through design space exploration, we configure two versions of the hardware accelerator and evaluate their performance on the Alveo U50 platform. Experimental results show that both configurations achieve 99.63 % bandwidth utilization. Compared to the CPU, the hardware accelerator achieves a 114.05× speedup and a 282.75× improvement in energy efficiency. It also outperforms the PyTorch implementation on the GPU, achieving a 29.81 ×speedup and a 297.87 ×improvement in energy efficiency. Additionally, the hardware implementation shows a 1.94 ×speedup and a 35.89 ×improvement in energy efficiency over the official CUDA-accelerated version.

Read the paper · More papers on PaperTik