Adaptive Sparse Deep Neural Network Inference on Resource-Constrained Cost-Efficient GPUs
Ming Dun, Xu Zhang, Huawei Cao, Yuan Zhang, Junying Huang, Xiaochun Ye · 2023
Sparse Deep Neural Networks (SpDNNs) has gained great popularity and been widely applied in various machine learning area. Compared to traditional dense DNNs, the unpre-dictable irregularity and sparsity in the sparse weight matrices of SpDNNs make them difficult to be efficiently parallelled. Moreover, most of the recent advanced efforts to optimize SpDNNs are based on high-end GPUs like NVIDIA V100, which may not be affordable to individuals and smaller re-search groups. However, migrating the SpDNNs to those cost-efficient but resource-constrained GPUs confronts enormous challenges, including limitations in both memory and computing resources, as well as the tiresome hyper-parameter tuning in batch parallelism. In this paper, we accelerate SpDNNs on GPUs with more restricted resources through exploiting the memory and computing resources. On one hand, we design the adaptive memory-aware data partition scheme to reduce memory consumption automatically. On the other hand, we propose the Tensor core/CUDA core fusion mechanism to efficiently utilize the hetergeneous computing resources on modern GPU architecture. To the best of our knowledge, we are the first to improve the performance on SpDNNs through adaptive memory tuning and utilizing hetergeneous computing core concurrency. We compare our implementation with the state-of-art previous champions, and the results demonstrate that our work achieves the highest speedup of 1.35 x and 1.39x compared to 2022 champion S&Z on single and multiple NVIDIA Tesla T4 GPUs, respectively. What's more, our work can reach similar or even better throughput compared to 2020 champion H&P on 6 V100 GPUs with only 4 T4 GPUs.