Hybrid Pre-processing Technique for Handling Imbalanced Data and Detecting Outliers for KNN Classifier

Preeti Nair, Indu Kashyap · 2019

Data mining is a technique of examining huge quanta of pre-existing data in order to discover new patterns and relationships among them, which will help to make better decisions. Classification is a data mining technique which organises data into categories. In this paper, in order to enhance the performance of k nearest neighbour (kNN) classifier-a kind of classification technique that is among the most widely used-a new data pre-processing technique has been proposed, which can handle some classification issues such as imbalanced data and outliers. In an imbalanced dataset, the classification categories are not equally distributed. Imbalanced dataset have an inherent issue when it comes to using classifiers on them that have been developed using machine learning algorithms. The basic nature of these algorithms is to reduce errors without relying on balance of classes. Another issue addressed in this paper is the matter of outliers or extreme values. Outlier or extreme values are those values that are outside the expected range of values. The quality of Classification modelling can be greatly enhanced by identifying and excision of these values. In this proposed technique two data pre-processing techniques have been combined to form a hybrid pre-processing technique. The two data pre-processing techniques are resample technique and inter quartile range technique (IQR). Some Imbalanced dataset with outliers that can be considered as yardsticks were taken for this study. It was observed that the classification results obtained were far superior to the classification done without the pre-processing technique.

Read the paper · More papers on PaperTik