GPU-Based Performance Analysis of Deep Learning Training

Mihiri Elapatha, Lahiru Wijethunga, Jananie Jarachanthan · 2024

Deep learning has been promising in recent years which learns complicated patterns and constructs data-driven decisions. To improve the effectiveness, due to the computational and memory requirements, most of the deep learning training workloads utilize GPUs. The challenges such as growing data sets, model size, and computation specifically count on the factors of the deep neural network used to train and the resource factors of GPUs utilized for that training. This study investigates those factors and represents the flow of them with the shifts. Since everything with deep learning training focuses the performance, this research designed and implemented a predictive model to find the end-to-end training time depending on the given resources by the user. The factors evaluated are model-specific parameters in-cluding batch size, matrix size, and kernel dimensions in addition to variables like GPU memory bandwidth, core count, and clock speed. This research offers practical insights for deep learning practitioners by benchmarking GPUs, such as NVIDIA's Tesla P100, K80, and M40, across several neural network architectures, including AlexNet, GoogLeNet, and VGG16. Furthermore, by creating predictive models for deep neural network training time on particular GPU setups, the research attempts to close the gap between theoretical expectations and real-world performance by guidance for cost-effective GPU selection, and optimizing resource allocation in deep learning processes. The analysis of the results for the different models in different GPUs presents ~5-20% factor with mitigation techniques to the users for the appropriate resource allocation.

Read the paper · More papers on PaperTik