Effects of Population Initialization on Evolutionary Techniques for Subgroup Discovery in High Dimensional Datasets
Vítor de Albuquerque Torreão, Renato Vimieiro · 2018
Many Evolutionary Algorithms have been proposed to solve the Subgroup Discovery task. Some of these, however, have been shown to work poorly in high dimensional problems. The best performing evolutionary algorithm for subgroup discovery in high dimensional datasets has a particular way to initialize its starting population, limiting the size of initial solutions to the lowest possible value. As with most population-based techniques, the outcome of evolutionary algorithms is usually dependent on the initial set of solutions, which are typically randomly generated. The impact of choosing one initialization technique over another in the final presented solution has been the topic of many published works in the broad area of evolutionary computation. However, to the best of our knowledge, it has not been the topic of study in the specific case of the Subgroup Discovery task, especially when considering high dimensional datasets. Therefore, this paper aims at studying whether or not it is possible to improve the performance of evolutionary algorithms in high dimensional subgroup discovery tasks by biasing the initial population to individuals with lower sizes.