On stochastic optimization for deep learning

Svetlana Volkova · 2024

Various stochastic optimization methods are utilized for training neural networks. The objective of this research is to provide a comprehensive overview of stochastic optimization methods proposed to enhance and expedite the convergence of neural networks. The article presents a highlighting of stochastic optimization methods' advantages and drawbacks, while also analyzing the constraints of their applicability. It commences with an introduction to the problem formulation, followed by sections dedicated to various algorithmic modifications: SGD-based stochastic optimization methods, adaptive gradient methods, and methods of adaptive moment estimation. In conclusion, the article underscores the importance of a judicious selection of a method, contingent upon its characteristics, applicability constraints, specific task, model architecture, and data quality.

Read the paper · More papers on PaperTik