A Case Study on Data Classification Approach Using K-Nearest Neighbor

Jogeswar Tripathy, Rasmita Dash, Binod Kumar Pattanayak · 2021 International Conference in Advances in Power, Signal, and Information Technology (APSIT) · 2021

Data mining is the process of obtaining knowledge and information from massive amounts of data. Data mining is mostly used for data analysis. In Data Mining various techniques are used that are association mining, regression, prediction, classification, clustering, etc. Classification is described as the process of identifying a collection of models (or functions) that explain and differentiate data classes and ideas, to use the model for detecting the classes of unknown objects or patterns, whose class designations aren't clear. Classification is a supervised learning problem. That means in machine learning, Classification is the problem of identifying data patterns from the group of patterns according to their characteristics and define which patterns are from which class. Pattern classification can be done by using various classifiers. A classifier is a program that inputs the feature vector of a pattern or data point and assigns it to one of a set of designated classes. The classifiers such as Artificial Neural Network (ANN), k-Nearest Neighbor (k-NN) classifier, Support Vector Machine (SVM), etc. are used for pattern classification purposes. Focusing on the classification technique of data mining, in this research work the accuracy of k-NN using three datasets from the UCI machine learning library is presented. The main goal of this paper is to provide a review to find out the accuracy of the k-NN classification technique using different datasets in data mining. The k-NN classifier is a simple but efficient approach used for classification in research.

Read the paper · More papers on PaperTik