Energy Efficient Systolic Array for Deep Learning Acceleration using Multi-Level Compression
M Sakthimohan, S Hanees, Elizabeth Rani G · 2025
The widespread use of artificial intelligence has reached to almost every application dealing with lot of real-world problems. The need for hardware and power efficient deep neural network (DNN) accelerators is very much important for computational and power-intensive nature of algorithms like DNNs. However, these accelerators have energy limitations when big DNNs are inferred in contexts with limited resources. Large DNNs require a significant amount of energy to obtain from the accelerator memory since they contain hundreds of millions of training parameters. A DNN based Multilevel Compress (DNNMLC), a low-power approach that applies conventional compression techniques relevant to hardware accelerators for real time applications with computational complexity, is suggested as a solution to this problem. The three-phase method includes hardware-based weight post-quantization trimming, dictionary and bitmask-based data optimization, and decompression by a low-complexity on chip architecture. Analysis has been done on the proposed DNNMLC and the results show that DNNMLC provides a maximum compression with an optimized memory usage in the processing element of DNN accelerator, without causing any performance degradation in big DNNs. Additionally, suggested low-power decoder utilizes a small area, making it possible to use DNNMLC in settings with limited resources. The proposed multi-level compression results in 45.83% compression efficiency compared to the conventional uncompressed DNN accelerator.