Tessellation-based multi-block memory mapping scheme for high-level synthesis with FPGA
auJuan Escobedo, auMingjie Lin · 2016
For many intensive computing tasks, simultaneous data access into multi-dimensional data arrays is highly restricted by its data mapping strategy and memory port constraint. As such, to increase memory accessing bandwidth, innovative memory partitioning and mapping algorithms have been proposed to simultaneously access multiple memory blocks through physically distributing data elements in the same logical array onto multiple memory blocks. Fortunately, FPGA device provides an unique opportunity of implementing application-specific memory infrastructure that maximizes memory access performance. However, even with the help of existing high-level synthesis (HLS) tools, customizing memory architecture still poses severe challenges that impede the performance of data path. In fact, existing memory partitioning and mapping schemes exploit either linear skewing or hyper-plane partitioning, therefore causing excessive run-time delay and non-optimal memory block space utilization. This work presents a hardware-efficient memory partitioning and mapping scheme with both low computing complexity and low hardware overhead for accessing multidimensional arrays. Targeting at affine memory access patterns often found in many data-intensive applications, our key idea is to leverage the geometric concept of tessellation widely known in combinatorial study and adopt a partitioning scheme based on geometric arguments instead of counting integer points in polytopes for intra-block offset generation. Aiming to assist HLS, our tessellation-based memory scheme exploits hidden memory parallelism through leveraging physically independent memory blocks in FPGAs. Using FPGA devices, our experimental results have shown that our memory partitioning algorithm saves up to 63.7% in the amount of arithmetic operations, around 15% in execution time, and 31.1% in storage overhead relative to the state-of-the-art approach on average across five widely-used circuit benchmarks.