Characteristics of Accuracy Function on Multiclass Classification Based on Best, Average, and Worst (BAW) Subset of Random Forest Model

Ronny Susetyoko, Wiratmoko Yuwono, Elly Purwantini, Budi Nur Iman · 2022 International Electronics Symposium (IES) · 2022

This study aims to determine the effect of percentage of training data and number of features on average accuracy of classification model. Characteristics of the accuracy function of 2 datasets with different feature data types were also studied. All possible random forest models were calculated for accuracy based on some percentage level of training data and number of features. The accuracy function was constructed based on subsets of best accuracy (Best), average accuracy (Average), and worst accuracy (Worst) using polynomial regression. There were no significant difference in the average accuracy of various percentage levels of training data. However, there were significant difference in the average accuracy based on the number of features. The best polynomial model for average accuracy function of the 2 datasets had R299.50% and R296.99%, respectively. Based on individual testing (with a 95% confidence level), there was a linear effect of the percentage of training data on the average accuracy, and there were linear and quadratic effects on the number of features on the average accuracy. In the first dataset, classification model using 7 features obtain classification accuracies ranging from 96.9% – 97.6%. In the second dataset using 7 features, the classification accuracies range from 74.0% – 75.0%. Characteristics of the Best, Average, and Worst accuracy (BAW) functions could be an important description of the accuracy function characteristics and to determine the number of features.

Read the paper · More papers on PaperTik