ML-Fusion: Determining Memory Levels for Data Reuse Between DNN Layers

Zikang Zhou, Xuyang Duan, Kaiqi Chen, Y Chen, Jun Shu Han · 2024

With the increasing complexity of applications and the improvement of computational power, modern neural networks (DNNs) have become more memory-intensive. To address the bandwidth problem, modern hardware architectures often incorporate multi-level memory to efficiently reuse data with different reuse distances. Additionally, a promising technique to reduce DNN bandwidth requirements is layer fusion which reuses inter-layer data at on-chip memory. However, previous studies on inter-layer data reuse scheduling have focused primarily on reusing at the outermost on-chip memory, neglecting the exploration of multi-level architecture, which represents a significant optimization space.

Read the paper · More papers on PaperTik