Scaling a Convolutional Neural Network for Classification of Adjective Noun Pairs with TensorFlow on GPU Clusters

Victor A. Campos, Francesc Sastre, Maurici Yagües, Jordi Torres, Xavier Giró-i-Nieto · 2017

Deep neural networks have gained popularity inrecent years, obtaining outstanding results in a wide range ofapplications such as computer vision in both academia andmultiple industry areas. The progress made in recent years cannotbe understood without taking into account the technologicaladvancements seen in key domains such as High PerformanceComputing, more specifically in the Graphic Processing Unit(GPU) domain. These kind of deep neural networks need massiveamounts of data to effectively train the millions of parametersthey contain, and this training can take up to days or weeksdepending on the computer hardware we are using. In thiswork, we present how the training of a deep neural networkcan be parallelized on a distributed GPU cluster. The effect ofdistributing the training process is addressed from two differentpoints of view. First, the scalability of the task and its performancein the distributed setting are analyzed. Second, the impact ofdistributed training methods on the training times and finalaccuracy of the models is studied. We used TensorFlow on top ofthe GPU cluster of servers with 2 K80 GPU cards, at BarcelonaSupercomputing Center (BSC). The results show an improvementfor both focused areas. On one hand, the experiments showpromising results in order to train a neural network faster. The training time is decreased from 106 hours to 16 hoursin our experiments. On the other hand we can observe howincreasing the numbers of GPUs in one node rises the throughput, images per second, in a near-linear way. Morever an additionaldistributed speedup of 10.3 is achieved with 16 nodes taking asbaseline the speedup of one node.

Read the paper · More papers on PaperTik