Optimizing the Hyper-parameters of Multi-layer Perceptron with Greedy Search

Mingyu Bae · American Journal of Computer Science and Technology · 2021

The core of deep learning network is hyper-parameters which are updated through learning process with samples. Whenever a sample is fed into deep learning network, parameters change according to gradient value. At this point, the number of samples and the amount of learning are crucial, which are batch size and learning rate. To find the optimal batch size and learning rate, lots of trial is inevitable so it takes so much time and effort. Therefore, there have been lots of papers to enhance the efficiency of its optimization process by automatically tuning the single parameter. However, global optimization can’t be guaranteed by simply combining separately optimized parameters. This paper propose brand new effective method for hyperparameter optimization in which greedy search is adopted to find the optimal batch size and learning rate. In experiment with Fashion MNIST and Kuzushiji MNIST dataset, the proposed algorithm shows the similar performance as compared to complete search, which means the proposed algorithm can be a potential alternative to complete search.

Read the paper · More papers on PaperTik