Multivariate LSTM for Execution Time Prediction in HPC for Distributed Deep Learning Training
Tasnim Assali, Zayneb Trabelsi Ayoub, Sofiane Ouni · 2024
In the last decade, Distributed deep learning has been widely used and introduced in research for highly computational tasks where time is very critical due to its capability to train deep learning models in less time within High-performance computing (HPC) clusters. However, Distributed Deep learning faces issues in heterogeneous HPC environments such as with synchronization that could slow down the training and affect the quality of training, especially in real-time applications. Therefore, in this article, we aim to improve the quality of training the distributed deep learning by using the multivariate LSTM model to predict the execution time of workers relying on its capability of capturing long-term dependencies in sequence prediction problems. Our LSTM model showed good results during training, validation, and testing with RMSE = 0.5284, with a dataset of a trace of training jobs running ML algorithms in Alibaba PAI. Also, we compared our model with CNN1D and GRU well-known models for regression problems and our model showed better results compared to them.