Parallel Training of a Back-Propagation Neural Network Using CUDA

Xavier Sierra-Canto, Francisco Madera-Ramírez, Víctor Uc-Cetina · 2010

The Artificial Neural Networks (ANN) training represents a time-consuming process in machine learning systems. In this work we provide an implementation of the back-propagation algorithm on CUDA, a parallel computing architecture developed by NVIDIA. Using CUBLAS, a CUDA implementation of the Basic Linear Algebra Subprograms library (BLAS), the process is simplified, however, the use of kernels was necessary since CUBLAS does not have all the required operations. The implementation was tested with two standard benchmark data sets and the results show that the parallel training algorithm runs 63 times faster than its sequential version.

Read the paper · More papers on PaperTik