Fast Approximate LUT-based Vector Multiplication in DRAM
Ziming Chen, Quan Deng, Yongwen Wang · 2023
Vector multiplication is widely used in real-world applications. To accelerate vector multiplication, processing-in-memory-based domain-specific architectures leverage lookup tables (LUTs) to decrease the computational complexity of multi-plication. However, the overhead of LUTs increases exponentially with the result space, which throttles the system performance.To decouple the LUT size and the performance improvement, we make a trade-off between accuracy and performance. We propose fast approximate LUT-based vector multiplication in DRAM, which builds the partial LUT of higher-bit locations and reuses the LUT to calculate the lower-bit data. We propose an operand reorganizing optimization to group more zero, which does not affect the results. To reduce the LUT and pre-calculation cost further, we propose a value encoding. Our experiment shows that the performance of the proposed design can be improved by 4x and 1.25x compared with LAcc in AlexNet and MobileNetV2 without any accuracy loss, respectively. Besides, the performance of the proposed design improves by up to 3.74x compared with pLUTo in multi-vector multiplication. The area and power of the proposed design are 54.8mm2and 5.35W, respectively.