Evolutionary cross validation

Thineswaran Gunasegaran, Yu–N Cheah · 2017

Cross validation methods such as 10 fold cross validation is used to generate folds randomly. A fold refers to the combination of training data subset and test data subset splits for training and validating machine learning models. Every fold would yield a certain accuracy value for the model. In the case of 10 fold cross validation, the overall accuracy is estimated by averaging the accuracy values produced by all 10 folds. For any dataset which contains a certain number of instances, there are many possible combinations of train-test data splits that could be generated. This implies that the train-test splitting process in cross validation can be formulated as a combinatorial optimization problem. The set of all possible train-test splits for a dataset form an evolutionary search space. In existing works, genetic algorithm based feature selection formulates the problem of feature subset selection as a combinatorial optimization problem. The objective of those works is to identify the most optimal feature subset which yields the best model accuracy. Apart from feature subsets, train-test splits or folds also influence model accuracy. Using different folds for training a model would result in varying ranges of accuracy. Out of all possible folds in the search space, a specific set of folds could yield optimal accuracy just like in the case of genetic algorithm based feature subset selection. This work proposes an evolutionary cross validation algorithm for identifying optimal folds in a dataset to improve predictive modeling accuracy. Results of experimental evaluation on several benchmark datasets suggest that the proposed algorithm provides significant improvement against the baseline 10 fold cross validation.

Read the paper · More papers on PaperTik