Empirical Assessment of Performance Measures for Preprocessing Moments in Imbalanced Data Classification Problem

Paweł Szeszko, Magdalena Topczewska · Lecture notes in computer science · 2016

The article concerns the problem of imbalanced data classification, when classes, into which elements belong, are not equally represented. In the classification model building process cross-validation technique is one of the most popular to assess the efficacy of a classifier. While over-sampling methods are used to create new objects to obtain the balance between the number of objects in classes, inappropriate usage of the preprocessing moment has a direct impact on the achieved results. In most cases they are overestimated. To present and assess this phenomenon in this paper three preprocessing techniques (SMOTE, Safe-level SMOTE, SPIDER) and their modifications are used to make new elements of data sets to balance cardinalities of classes, and two classification methods (SVM, C4.5) are compared. k -folds cross-validation technique ( \(k=10\) ) considering two moments of preprocessing approaches is performed. The measures as precision, recall, F -measure and area under the ROC curve (AUC) are calculated and compared. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik