Comparison of the stochastic gradient descent based optimization techniques
Ersan YAZAN, Muhammed Fatih Talu · 2017 International Artificial Intelligence and Data Processing Symposium (IDAP) · 2017
The stochastic gradual descent method (SGD) is a popular optimization technique based on updating each θkparameter in the ∂J(θ)/∂θkdirection to minimize/maximize the J(θ) cost function. This technique is frequently used in current artificial learning methods such as convolutional learning and automatic encoders. In this study, five different approaches (Momentum, Adagrad, Adadelta, Rmsprop ve Adam) based on SDA used in updating the θ parameters were investigated. By selecting specific test functions, the advantages and disadvantages of each approach are compared with each other in terms of the number of oscillations, the parameter update rate and the minimum cost reached. The comparison results are shown graphically.