Characterize and Compare the Performance of Deep Learning Optimizers in Recurrent Neural Network Architectures
Mohammad Zaeed, Tanzima Zerin Islam, Vladimir Inđić · 2024
The emergence of newer hardware and limited accessibility to that hardware are driving the need to characterize application performance on commonly used architectures to gain valuable insights. While abundant performance events create opportunities for detailed performance characterization, they also make the process overwhelming. This paper takes a systematic approach to address the problem of performance characterization by taking advantage of such opportunities. Specifically, this paper 1) characterizes the hardware usage behaviors of commonly used optimizers in the context of Recurrent Neural Network (RNN) architectures, 2) connects observations to identify problems and opportunities for optimization in those optimizers, and 3) creates a large performance dataset. Comparing the hardware resource usage behaviors of Gradient Descent (GS), AdaGrad (AdaGrad), and ADAM (ADAM) optimizers for three commonly used recurrent neural network-based deep learning architectures running on AMD MI100 GPUs, this paper identifies non-uniform memory access as one of the most frequent issues and suggests strategies for optimizing such problems.