Evolutionary Swarming Particles To Speedup Neural Network Parametric Weights Updates

Subhayu Dutta, Subhrangshu Adhikary · 2023

Gradient-based neural networks weight optimizers such as stochastic gradient descent, adam, adagrad or adadelta provides robust and reliable results, and they are very slow to train as they slowly and steadily iterate over the solution search space. These days deep learning models are getting very large, having billions or even trillions of parameters. Updating these many parameters with gradient-based methods would require large computational resources and time. Therefore, an alternative, reliable and fast parametric update rule is required. Particle swarm optimizer is a popular algorithm used for solving various optimization problems like design control systems, feature selection, signal denoising, compression techniques, etc. In the state of the art, the closest application of particle swarm in a neural network is to tune hyperparameters to find the best combinations for a given set of problems. But no study has explicitly used evolutionary swarming particles to directly update weights and biases of a neural network. With some modifications on both the general batch weight update mechanism of a simple neural network and the particle swarm, we successfully combined the two together. The tests were performed on two datasets, one for kidney stone prediction and one for breast cancer classification showing that the classification accuracy with the proposed modification is better than the state of the art in multiple cases. But the most important observation was that the proposed model was found to be up to 147 times faster than AdaDelta and 110 times faster than stochastic gradient descent.

Read the paper · More papers on PaperTik