Tuning Distance Metrics and K to Find Sub-categories of Minority Class from Imbalance Data Using K Nearest Neighbours
Md Saidur Rashid Mahin, Md. Jahidul Islam, Biplab Chandra Debnath, Ayesha Khatun · 2019
Data classification by the machine learning algorithms is one of the most studied topic within data mining. Machine learning classifiers suffers significantly while learning from the imbalance data. Within imbalance class data, a class covers a very small portion of the whole dataset also known as minority class or the positive class, generally are the most significant. The issue that plays the key responsibility for these degradation's of performance is primarily the distribution of different class of samples within the dataset. It is the presence of local sub-concepts from different classes within the domain of another class. Recent time, a number of studies have concentrated on categorizing the minority class into several sub-categories or sub-concepts based on their logical neighbourhood. For this purpose-these studies have utilized methods like K nearest neighbour and kernel method. This study aimed at improving the categorization process of the minority class by incorporating an idea of using dataset specific distance function for the categorization process. The major objective of this study is to choose a distance metric among five distance metrics and a k value for a specific dataset, in which the classification performance is most optimal. For the analysis purpose five datasets are chosen.