A Many-Objective optimization Approach for Complexity-based Data set Generation
Thiago R. Fraca, Pericles B. C. de Miranda, Ricardo B. C. Prudêncio, Ana C. Lorenaz, André Câmara Alves do Nascimento · 2020
The assessment of machine learning algorithms in a particular task is usually done by means of empirical evaluation on real world observational data. However, sometimes there is no previously annotated data available. Synthetic datasets have gained attention as an alternative for efficient classifier evaluation, since they are accessible and high parameterizable for a given learning task. The characterization of such databases can be done by means of descriptors on the learning object at hand, e.g., complexity measures, which extract statistical and geometric characteristics from the data sets, for a given classification problem, in order to estimate their complexity. Such complexity measures can be used to guide the production of synthetic datasets, so it reinforces one or more dataset features. The present work proposes the use of a many-objective algorithm for the generation of synthetic data considering four measures of complexity that will be balanced at the same time. The results showed that the proposal is able to optimize conflicting objectives, generating datasets of specific complexities.