Hybrid Multi-tile Vector Systolic Architecture for Accelerating Convolution on FPGAs

Jay Shah, Nanditha Rao · 2024

To enhance the efficiency of image-kernel convolution operations in convolutional neural networks, we introduce a Vector Systolic Array Accelerator with adaptable lane-width. This architecture utilizes multiple vector lanes for concurrent data element processing to address computational bottlenecks. To further improve the throughput, we propose a novel hybrid tiled vector systolic design in which we partition the hardware resources to efficiently utilize them along with a unique data mapping strategy. In this approach, we choose some tiles to use LUTs and others to use DSPs. We observe that the throughput of the vector systolic accelerator is 7x and 1.26x higher for single-tile and multi-tile configurations than their non-vector counterparts respectively. The hybrid tile design significantly increases tile count, achieving a competitive peak throughput of 1165 GOPs and 1072 GOPs for optimal lane width of Vector-6 and 8, which is 3.8x better than related work.

Read the paper · More papers on PaperTik