Memory Footprint Optimization for Neural Network Inference in Mobile SoCs
Sakshi Agrawal, Priyankar Ghosh, Gaurav Kumar, Tripuraneni Radhika · 2023
In the last decade, deep neural networks (DNNs) have gained widespread acceptance and used heavily in mobile SoCs for performing complex tasks. Memory footprint is one of the key factor that determines the overall performance of the mobile platform while running DNN inference because DNNs use significant amount of memory for storing the weights and the intermediate feature maps computed by the hidden layers. Since the available DRAM in mobile SoCs is limited, memory footprint optimization becomes a key challenge. In this paper we address the problem by proposing two strategies for optimizing the DRAM allocation. The first strategy targets to increase the utilization of the different buffers used during the inference and uses best fit iteratively to come up with a better allocation. In the second strategy we explore a variant of the best-fit to address the allocation of feature maps related to concat operation. Experimental result on a set of proprietary DNNs show that the proposed strategy could reduce memory footprint by 15% on an average. For DNNs of certain category, more than 35% reduction is observed.