A Memory Management Method for DNN Inference for Aerospace Embedded Systems
Shen Li, Fei–Yue Wang, Shugang Zhang, Tianxing Wang, Zhong Ma, Feng Liang · 2024
Currently, deep neural networks(DNN) are widely used in aerospace field, however, due to the embedded system in aerospace field, it is often difficult to support large-scale in-orbit inference calculation of deep neural networks due to the resource constraints in memory and power consumption. In order to effectively reduce the memory space occupied by deep neural networks during inference computation, we propose a memory management method for deep neural network inference computation by establishing a two-dimensional model of size and lifetime for memory allocation, and by fusing three strategies of time-first, size-first, and area-first in the consideration of both depth traversal and breadth traversal of the computational graph. In order to reduce the memory consumption of intermediate computation results by deep neural networks during inference computation, we determine the optimal memory reuse scheme between different intermediate computation result tensors by arranging the intermediate computation result tensors in the two-dimensional plane without overlapping. We also tested typical deep neural network algorithms based on the self-developed TIANJI NPU, and the test results show that our method has achieved a large improvement compared with the original method of TIANJI NPU software toolchain, and the maximum reduction in memory usage reaches 85.51%, and the minimum reduction is 21.31%.