HiSo: Co-optimizing the Intra-layer and Inter-layer Scheduling Schemes with the Hybrid Data Flow for PIM Architectures

Yunping Zhao, Sheng Ma, Jianmin Zhang, Tiejun Li, Yuhua Tang · ACM Transactions on Architecture and Code Optimization · 2025

The Processing-In-Memory (PIM) architecture becomes a promising candidate for realizing energy-efficient CNN acceleration by integrating computation and memory. For deploying a CNN, the PIM-based designs need a scheduling scheme to translate massive hardware resources into actual performance. The scheduling scheme includes the intra-layer scheduling scheme to map data of one layer and the inter-layer scheduling scheme to allocate resources for multiple layers. Currently, the research of scheduling schemes for PIM-based studies is in its early stages and faces the following limitations. First, the intra-layer scheduling scheme mainly uses the Weight Stationary (WS) data flow to avoid weight movements. However, the amount of activation movement is larger than that of the weight movement for some DNNs or layers, making the activation movement become the bottleneck for reducing latency and energy consumption. Second, the Input Stationary (IS) data flow provides new opportunities to further reduce activation movements. However, the traditional IS data flow introduces a huge amount of across-crossbar computations and relies on the 3D-ReRAM structure. Third, the layer-wise processing structure of the DNN introduces complex inter-layer data dependence and imbalance workload, decreasing pipeline efficiency. Fourth, there is little study on co-optimization of intra-layer and inter-layer scheduling schemes. Inspired by these observations, we propose the Hybrid data flow based Intra-layer and inter-layer Scheduling schemes Optimization framework for PIM-based architectures, named HiSo. Specifically, we propose a novel data flow, named PP-IS, that replaces the convolution unrolling with selecting activations according to convolution windows to reduce activation movements. We also define and optimize the partition methods to improve data reuse and computational parallelism for the WS and PP-IS data flows. Then, we trade off performance and energy efficiency by using the hybrid data flows. Finally, we further improve energy efficiency and performance by co-optimization of intra-layer and inter-layer scheduling schemes. Our experimental results and analysis demonstrate the potential of the HiSo. Compared to the state-of-the-art design, the NEBULA, the HiSo can significantly improve energy efficiency, performance, and power efficiency by 1.03× ∼ 17.98×, 1.6× ∼ 50.7×, and 1.2× ∼ 229×, respectively.

Read the paper · More papers on PaperTik