A chain-multiplier for large scale matrix multiplication

Can Wei, Yukun Song, Duoli Zhang · 2017

Matrix operation has high time complexity and traditional serial algorithm is less efficient. In the existing design of matrix multiplication, systolic array and other methods are usually used for hardware acceleration. But as the scale of matrix computing increases, the “storage wall” problem caused by data throughput bandwidth has become the bottleneck of performance improvement. In this paper, a hardware accelerator suitable for large scale matrix multiplication is proposed. This design takes non-buffering organization, the data entered into the computing unit directly involved in the operation, reducing the storage pressure. In addition, it has low demand for data throughput, avoiding the impact of “storage wall”, while maintaining high operational performance. The multiplier also supports on-line configuration of operation scale, which is suitable for matrix operation of different scales. The design was performed on XC7V2000T chip for prototype verification, which can integrate up to 1080 processing elements. After testing, when the processing elements number is 256, the peak bandwidth is 19.2 Gbit/s, and in operation 1k-order matrix, compared with the same processing element number of systolic structure, our design compressed 1k times bandwidth and achieved the same operational performance.

Read the paper · More papers on PaperTik