Training Convolutional Neural Networks with Differential Evolution using Concurrent Task Apportioning on Hybrid CPU-GPU Architectures

Rochan Avlur Venkat, Zakaria Oussalem, Arya K. Bhattacharya · 2021

The core algorithm for training of Artificial Neural Nets (ANNs) continues to remain the back-propagation (BP) algorithm - to an extent that it is now considered as a paradigm of Deep Learning (DL). Many important facets of DL, like hierarchical construction of features across layers in image recognition, vanishing gradients, etc., are taken for granted without recognizing that these may implicitly be induced by BP itself. Evolutionary Algorithms (EAs) perform global optimization in contrast to localized gradient descent of BP. If used extensively for ANN training, they can potentially disrupt these assumed facets of DL - and construct alternative and interesting perspectives. But they are severely constrained by the need for large computational resources, as they work concurrently on a population of candidate solutions. The bulk of processing occurs in the forward pass through the ANN of thousands of data samples - which can be efficiently parallelized on GPUs. However, the candidates themselves can be launched in small groups on different CPU cores - by exploiting their natural concurrency. Here we explore the possibility of launching training of ANNs with EAs on hybrid CPU-GPU systems with candidates split across CPUs and samples across GPU threads. We conduct a series of experiments from which we synthesize a successful mechanism for orders-of-magnitude speedup of EAs through efficient apportioning of multi-level computing tasks onto different classes of processing elements. This enables analysis of DL using EAs empowering alternative interpretations of the above facets.

Read the paper · More papers on PaperTik