High Throughput Matrix Transposition on HBM-Enabled FPGAs
Yang Yang, Rajgopal Kannan, Viktor K. Prasanna · 2025
Matrix transposition is a classic operation in machine learning and scientific applications. HBM-enabled FPGAs, with their high-bandwidth capabilities, are increasingly deployed in the data centers. However, achieving high bandwidth utilization on HBM is challenging due to the large strided access patterns in matrix transposition, which significantly degrade bandwidth utilization. Additionally, saturating HBM bandwidth requires a large number of parallel accesses, further complicating the design. In this paper, we present a high throughput matrix transposition design for HBM-enabled FPGAs. Our design performs small strided accesses to HBM, ensuring optimized bandwidth utilization. We use on-chip SRAMs to reorganize data from HBM into the access pattern needed for matrix transposition. Inspired by Latin Squares, we propose a novel data layout for storing the matrix tiles in SRAM. This data layout is paired with a customized scheduling strategy to eliminate SRAM bank conflicts. We develop a fully pipelined architecture with multiple Processing Elements (PEs) to enable parallel HBM accesses. Our design is highly scalable, supporting various configurations of HBM channels and arbitrary matrix dimensions. We implement the proposed design on the AMD Alveo U280 FPGA. Experimental results show that our design achieves a matrix transposition throughput of up to 415 GB/s, more than 90% of the peak HBM bandwidth of the target FPGA platform. Our design outperforms state-of-the-art GPU implementations, delivering up to 1.44 × higher HBM memory bandwidth utilization.