Visualising Adam Oscillations in Neural Network Loss Landscapes
Henri van der Grijp, Anna Sergeevna Bosman, Katherine Mary Malan · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2025
The Adam optimiser is a popular adaptive gradient-based learning algorithm used to train neural networks. Adam uses adaptive learning rates (estimated per neural network weight) to increase the speed of convergence. While Adam is claimed to automatically perform step size annealing, it is not immune to oscillations, i.e., updating weights in inconsistent directions. This paper investigates the oscillatory behaviour of Adam in the context of computer vision classification tasks under different batch sizes. We confirm that Adam is indeed highly susceptible to oscillations. Further, we study starting gradient and magnitude of randomly initialised weights as predictors of weight saliency. Conversely to common belief, we discover that the initial gradient is not a reliable predictor of weight saliency. We attribute this finding to inherent oscillations. We also discover that larger batch sizes are more likely to yield inactive weights, bearing important implications for weight pruning.