COMPARISON OF VARIOUS CLASSIFICATION ALGORITHMS ON IRIS DATASETS USING WEKA

Kanu Patel, Jay Vala, Jaymit Pandya · International Journal of Advance Engineering and Research Development · 2014

Classification is one of the most important task of data mining. Main task of data mining is data analysis. For study purpose various algorithm available for classification like decision tree, Navie Bayes, Back propagation, Neural Network, Artificial Neural, Multi-layer perception, Multi class classification, Support vector Machine, k-nearest neighbor etc. In this paper we introduce four algorithms from them. Study purpose we take iris.arff dataset. Implement this all algorithm in iris dataset and compare TP-rate, Fp-rate, Precision, Recall and ROC Curve parameter. Weka is inbuilt tools for data mining. So we used weka for implementation. Keyword: classification, K-nn, ROC, FP-rate, Decision tree, WEKA I. INTRODUCTION Generally, data mining (sometimes called data or knowledge discovery) is the process of analyzing data from different perspectives and summarizing it into useful information - information that can be used to increase revenue, cuts costs, or both. Data mining algorithms which carry out the assigning of objects into related classes are called classifiers. Classification algorithms include two main phases; in the first phase they try to find a model for the class attribute as a function of other variables of the datasets, and in the second phase, they apply previously designed model on the new and unseen datasets for determining the related class of each record (1)(3). There are different methods for data classification such as Decision Trees (DT), Rule Based Methods, Logistic Regression (LogR), Linear Regression (LR), Naive Bayes (NB), Support Vector Machine (SVM), k-Nearest Neighbor (k-NN), Artificial Neural Networks (ANN), Linear Classifier (LC) and so forth. The comparison of the classifiers and using the most predictive classifier is very important. Each of the classification methods shows different efficacy and accuracy based on the kind of dataset (2) . In addition, there are various evaluation metrics for comparing the classification methods that each of them could be useful depending on the kind of the problem. Among the other criteria for comparing the classification methods, one could mention; precision, recall, error rate, confusion matrix. In this article, using a new method, five usual data classification methods (Decision tree, Multi- layer perception, Naive Bayes, C4.5, SVM) have been compared based on the AUC criterion. These mentioned methods have been applied on the random generated datasets which are

Read the paper · More papers on PaperTik