Implementation of Parallel Optimization Algorithms for NLP: Mini-batch SGD, SGD with Momentum, AdaGrad Adam
Wendi Huang · Applied and Computational Engineering · 2024
With the rapid development of machine learning technology, optimization algorithms and optimizers have become key to the development of related technologies contemporarily. Models need the help of optimizers to meet other performance indicators while saving computing resources. This research focuses on comparisons between optimizers, in the context of text sentiment classification tasks. The optimizers mainly compared in this article are mini batch SGD, momentum SGD, Adagrad and Adam. Through comparative experiments, it was found that SGD and its variants have a high dependence on the initial learning rate setting, while the performance of Adagrad and Adam is relatively balanced. Although the training time of Adagrad is shorter than that of Adam, its principal formula has flaws, which are not reflected in this task. The conclusions drawn in this article through comparison can point out the advantages and disadvantages of each optimizer, and can help realize better optimizers in subsequent research.