Learning Data Space Transformation Matrix from Pruned Imbalanced Datasets for Nearest Neighbor Classification

Seba Susan, Amitesh Kumar · 2019

The nearest neighbor classifier is deemed to be the litmus test for the worst-case scenario and has been often been relied on by data mining researchers to test the robustness of their algorithms. However not all datasets are suited for the distance-based classification, and a prior transformation of the data space generally helps. This paper proposes a novel hybrid sampling with data space transformation that boosts the performance of the nearest neighbor classifier for imbalanced datasets. The SSOMaj-SMOTE-SSOMin three-step pruning and resampling technique, introduced in a recent work by the authors, is used in the first stage of our experiments for achieving a balance between the under-represented minority and over-represented majority class. In the second stage, the transformation matrix learnt from the pruned dataset is used to transform the pruned training distribution and the test sample space. The proposed method thus enforces the spatial arrangement of the sampled training dataset into the test sample space. A variety of popular data space transformation techniques are investigated for the application. The consequences of transforming the original training space, based on the transformation learnt from the pruned training set are also investigated. Experiments on benchmark datasets with comparison to the state-of-the-art prove the supremacy of our learning approach for imbalanced datasets.

Read the paper · More papers on PaperTik