Evaluation of Stratified Validation in Neural Network Training with Imbalanced Data
Tomoumi Takase, Satoshi Oyama, Masahito Kurihara · 2019
Validation data is widely used for measuring the generalization performance of models in machine learning such as neural network training. However, because usually it is randomly extracted from training data, imbalanced validation data can be obtained, resulting in an unstable validation for models. To deal with this problem, we have used a stratified validation, which is a conventional method for preparing balanced validation data by making the ratio of the numbers for each class of the validation samples the same as that of the training samples. We have evaluated the stratified validation for a multilayer perceptron and convolutional neural networks with several benchmark datasets.