Supporting Moderate Data Dependency, Position Dependency, and Divergence in PIM-based Accelerators
Marzieh Lenjani, Kevin Skadron · IEEE Micro · 2021
Processing in memory (PIM) can alleviate the data movement overhead. However, PIM units built inside memory layers have low frequency, and this requires high parallelism to compensate for the low clock frequency. Single-instruction–multiple-data (SIMD) architectures can provide high parallelism for PIM with low overhead per arithmetic logic unit (ALU) operation. In SIMD, multiple ALUs perform the same instruction, and the processing unit accesses multiple consecutive words at once. Therefore, the control and access overhead is amortized among ALU operations. However, SIMD units cannot fully exploit the available word-level parallelism for 1) operations with data/position dependency or 2) operations with divergence (where different operations are performed on different words). A recent work, Fulcrum, proposes a subarray-level PIM design with high parallelism. This article discusses how Fulcrum alleviates the control and access overhead while exploiting word-level parallelism for operations with data/position dependency and divergence. We evaluate Fulcrum against bank-level SIMD architectures to highlight these benefits.