Using Complexity Measures to Evolve Synthetic Classification Datasets
Vinícius Veloso de Melo, Ana Carolina Lorena · 2018
Machine Learning studies usually involve a large volume of experimental work. For instance, any new technique or solution to a classification problem has to be evaluated concerning the predictive performance achieved in many datasets. In order to evaluate the robustness of the algorithm face to different class distributions, it would be interesting to choose a set of datasets that spans different levels of classification difficulty. In this paper, we present a method to generate synthetic classification datasets with varying complexity levels. The idea is to greedly exchange the labeling of a set of synthetically generated points in order to reach a given level of classification complexity, which is assessed by measures that estimate the difficulty of a classification problem based on the geometrical distribution of the data.