A Regression-based Model for End-to-End Latency Prediction for DNN Execution on GPUs

Ying Li, Yifan Sun, Adwait Jog · 2023

Deep neural networks (DNNs) have become increasingly popular in many domains as they reduce the requirement for human effort. However, today’s DNN applications suffer from high computational complexity and sub-optimal device utilization. To solve this problem, researchers have been proposing new system design solutions, which require performance models to help them with pre-product concept validation. This paper discusses how to build a simple, yet accurate, performance model for DNNs on GPUs. Our observations demonstrate prevalent linear relationships between the GPU execution times and operation counts of DNNs layers. Our proposed linear-regression-based execution time predictor can make predictions with an error rate of 28%.11This material is based upon work supported in part by the Google Research Scholar Award and William & Mary. This work was performed in part using the computing facilities at William & Mary and Google Cloud. This work was done while Jog was with William & Mary. Jog is currently with the University of Virginia.

Read the paper · More papers on PaperTik