Adaptive Condensed Nearest Neighbor for Imbalance Data Classification

Nijaguna Gollara Siddappa, Thippeswamy Kampalappa · International journal of intelligent engineering and systems · 2019

Classification over numerous real-world datasets has a peculiar drawback called unstructured class problem.A dataset is said to be unstructured when the majority of the class has more samples insignificantly than the minor class.Such drawbacks result in an ineffective performance of data classification techniques.Classification is a supervised learning method which acquires a training dataset to form its model for classifying unseen examples.However, in unstructured data classification, the class boundary learned by standard machine learning algorithms can be severely skewed toward the target class.As a result, the false-negative rate can be excessively high.The researches focus on the unstructured data classification using uncertain Nearest Neighbor (NN) decision rule and also found the major issues face by k-Nearest Neighbor (k-NN).In any case, given a dataset, prediction of accuracy is a monotonous task to improve the execution of kNN by tuning 𝑘.Because of class imbalance, the performance of kNN decreases and this situation is represented by dissimilar characteristic from various classes.This paper addresses the issues faced by kNN by developing Adaptive-Condensed NN (Ada-CNN).The Ada-CNN classifier utilizes the distribution and density of test point's neighborhood and learn an appropriate point-explicit 𝑘 by using artificial neural systems.Ada-CNN performed well compared to kNN and other well-known classifiers.The experimental results showed that Ada-CNN achieved nearly 94% accuracy for Diabetes dataset and 100% accuracy in pop-failure compared to kNN for imbalanced classification.

Read the paper · More papers on PaperTik