HPA: A Hybrid Data Flow for PIM Architectures

Sheng Ma, Yunping Zhao, Yuhua Tang, Yi Dai · 2024

The Processing- In- Memory (PIM) architecture becomes a promising candidate for deep learning acceleration by integrating computation and memory. Due to the simple mapping method and high parallelism, the Weight Stationary (WS) data flow is widely used in PIM-based studies to improve performance and energy efficiency. However, the WS data flow leads to huge activation movements, becoming the bottleneck for reducing latency and energy consumption. To address this issue, the Input Stationary (IS) data flow stores activations instead of weights in the PIM architecture to reduce data movements. However, the traditional IS data flow also faces several challenges. First, the inter-layer data dependence and imbalance workload decrease pipeline efficiency. Second, the across-array computation reduces energy efficiency and performance. Third, the traditional IS data flow relies on the 3D ReRAM structure. Inspired by these observations, we propose a Hybrid data flow for PIM Architectures, named HPA. The HPA contains novel intra-layer and inter-layer data flows, named the PP-IS data flow and the IS- WS hybrid data flow, respectively. The PP- IS data flow optimizes the data mapping strategy and computing method to reduce activation movements. In addition, the PP- IS data flow uses the parallel computing method to decrease across-array computations. Based on the novel intra-layer data flow, we propose the IS- WS hybrid data flow to trade off performance and energy efficiency. Finally, we optimize the pipeline for the hybrid data flow to mitigate data depen-dence and balance inter-layer workloads, improving pipeline efficiency. Our experimental results and analysis demonstrate the potential of the HPA. The performance and power efficiency of the HPA reaches 1.64$GFLOPS\sim 63$G F LO P Sand 2.1$TOPS/W\sim 151\ TOPS/W$, respectively. Compared to the state-of-the-art design, the NEBULA, the HPA can significantly improve power efficiency and performance by$22.1\times$and$7.8\times$, respectively, when deploying the MobileNet VI.

Read the paper · More papers on PaperTik