Parallel approaches to training feedforward neural nets
Louis Coetzee, Elizabeth C. Botha, Etienne Barnard · 1996
Neural networks have gained prominence and are used for the successful solution of many real-world problems. However, training these networks is difficult and time consuming. In this thesis we investigate parallel neural-net training algorithms which reduce the required time to train large networks. We focus our attention on feedforward multi-layer perceptron neural nets with sigmoidal transfer functions which are used extensively in pattern recognition applications. We start with a theoretical analysis of block backpropagation, an established online training algorithm. With a theoretical analysis and experimental verification we show that block backpropagation improves on general backpropagation (the standard online training algorithm). Once it is established that block backpropagation is superior to backpropagation, we propose several parallel implementations suitable for tightly as well as loosely-coupled MIMD (multiple instruction multiple data) architectures. In conjunction with the parallel implementations, we propose speedup models for each implementation and architecture which are used to analyse the performance of the algorithms. We conclude that online block backpropagation can be successfully parallelised. We then present a parallel implementation of a conjugate-gradient training algorithm using shared-memory constructs. We implement two versions on a distributed shared-memory MIMD architecture. One version is implemented with native code, and the other with P4, a portable parallel user library. We propose a speedup model which we use to analyse and compare the different versions. Our experimental approach, combined with an analysis of the speedup model, show that both versions are successful. Finally, we present a coarse-grain implementation of a batch-mode conjugate-gradient algorithm, which uses PVM to combine a distributed network of work-stations into a virtual parallel architecture. We once again propose a model of speedup which we use in conjunction with our experimental approach to investigate the feasibility of distributed workstations as nodes of a parallel machine. We conclude that for problems consisting of large training sets, the use of distributed workstations is a viable alternative to dedicated MIMD parallel architectures. Evaluating all the experimental results we conclude that parallel neural-net training algorithms are a viable alternative to sequential training, with the performance gains outweighing the initial parallel implementation costs.