MLMap: A Multilevel Mapping Flow for Coarse Grained Reconfigurable Architecture

Jiali Yu, Weidong Yang, Weiguang Sheng · 2022 IEEE 5th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC) · 2022

Coarse-Grained Reconfigurable Architecture (CGRA) is an attractive embedded system that can achieve both high energy efficiency and flexibility. But with the advent of the big data era, the bandwidth bottleneck brought by the traditional memory packaging of CGRA has become more rigorous. The 3D CGRA uses 3D stacking technology, like Hybrid Memory Cube (HMC), to alleviate this problem. It has very high total bandwidth, but not high bandwidth per vault. However, the conventional CGRA's compilation method doesn't consider 3D CGRA's multiple vaults architecture, distributed memory and the trend of the increasing number of Processing Elements (PEs), and greatly hinders performance improvement. To address these problems, we draw into the idea of supernode, which is an intermediate layer between the data flow graph (DFG) node and the DFG itself. Supernodes are used to capture data flow information, exploit regularity and apply optimization. Based on that, we propose a novel compilation flow for 3D CGRA, which decomposes the mapping into intra-vault and inter-vault levels including supernode partition, supernode-aware mapping, virtual mapping and data-aware unrolling. It maps the supernodes and DFG nodes to CGRA hierarchically to maximize the data reuse in memory to save the bandwidth for each vault. Besides, we integrate partition and unrolling into our multilevel mapping flow to further process the supernodes, which reduces bank conflicts, the execution time of each supernode, and instruction switch cost. Compared with the state-of-the-art mapping technology transplanted to 3D CGRA, our multilevel mapping flow shows a better performance of about$3.30\times$on average.

Read the paper · More papers on PaperTik