HMO: Host Memory Optimization for Model Inference Acceleration on Edge Devices
Chaoxiong Yi, Songlei Jian, Yusong Tan, Yusen Zhang · 2024
Deep learning (DL) is characterized by its demanding computational and memory requirements, which creates a significant challenge when deploying on edge devices. These devices often have limited computational capabilities and constrained resources. Most existing methods primarily focus on model-level techniques, such as model pruning or parameter quantization, to reduce model size and computation for accelerating inference. Considering the prevalent programming paradigms in DL, we propose a host memory optimization method, namely HMO, which can be integrated into DL programming framework, e.g., PyTorch, to improve the inference efficiency of DL models without modifying any model code. We particularly focus on memory optimization for intermediate variables in inference, aiming to enhance inference speed while maintaining a lower memory footprint. HMO involves a single profiling of inference to gather memory statistics about intermediate variables. These statistics are then used to guide subsequent inference. Additionally, we incorporate huge pages in operating systems to improve the memory access performance of HMO. Our experimental results show that HMO can achieve an average inference latency optimization ratio of 20.13 % compared with native PyTorch on six typical DL image representation models while effectively managing memory usage. Importantly, this is achieved without compromising model accuracy.