Analyzing the impact of data representations in classification problems using clustering

Felipe Farias, Teresa B. Ludermir, Carmelo J. A. Bastos-Filho, Flávio Rosendo da Silva Oliveira · 2019

This work presents an investigation about how to better represent output data labels to be used in supervised training of classifiers. The posed hypothesis is that grouping cohesive patterns into clusters and assigning them sub-labels, may improve the classifier performance. We used 12 benchmark datasets to test our hypothesis. First, we create the clusters, and when appropriate, new sub-labels were generated, according to Fuzzy-CMeans and Silhouette score thresholds. After that, Multilayer Perceptrons were employed to model each dataset with cluster generated sub-labels, obtaining promising results. From results, we observed that in cases where the sub-labels were used, the accuracy increased with statistical significance with p=0.05 in 22 cases and remained statistically equivalent in 14 cases, presenting no decrease in accuracy.

Read the paper · More papers on PaperTik