Improving Inference Time of Deep Learning Model with Partial Skip of ReLU-fused Matrix Multiplication Operations
Sungkyun Kim, Jaemin Kim, Nahun Kim, Mincheal Kang, Jiwon Seo · 2022
Deep learning has been expanding its application, while large-scale models tend to perform well. However, as such a model inevitably requires a vast amount of resources and computations, lengthy inference time is a crucial, but essential, consequence that needs to be optimized for the efficient utilization of deep learning. To achieve the goal, we aim at fusing the Rectified Linear Unit and matrix multiplication in the inference process, which we may reduce the total amount of computation by predicting the sign bit of output value. We propose four methods of prediction and statistically choose an optimal method for reducing inference time with low accuracy loss.