Performance Evaluation of Distributed Training in Tensorflow 2
Nguyen Quang-Hung, Hieu Nguyen Doan, Nam Thoai · 2020
Deep learning (DL) is most interested in Artificial Intelligence (AI) because it brings many interesting results in object identification, natural language processing, speech processing, application algorithms in autonomous vehicles, etc. DL also use neutron network. A useful neural architecture for an application may have tens or hundreds of millions of numbers that need to be selected appropriately during the training process. The most popular technique for training this architecture is a gradient-based iteration technique. To complete the training process, it needs to iterate over large amounts of data, which is why the training process is so slow. Without high performance computing (HPC) technology, it will take weeks or months to complete the training of network architecture. Training DL by distributed training is one of the solutions to improve the learning process faster and reduce computing time. In this paper, we study performance evaluation for distributed training strategy of the Tensorflow 2.2 on some GPUs.