Filter Selection Methods for Multiclass Classification
Rhodessa J. Cascaro, Bobby D. Gerardo, Ruji P. Medina · 2019
Feature selection is used in choosing relevant features that contribute to the predictive power of a machine learning model. Irrelevant features tend to decrease model accuracy and cause overfitting. Feature selection is the solution to dimensionality problems, especially that data nowadays are unstructured. There are three types of feature selection techniques; filter, wrapper and embedded. Filter methods uses statistical scoring and ranks features by the score. Wrapper methods uses a model to select features and evaluates according to model accuracy. Embedded methods combine the properties of both previous algorithms and selects features in the process of training. This study is focused on the filter types, specifically Chi-square, Information Gain and Relief. On the other hand, multiclass classification is a task that involves classifying instances into three or more classes. This paper aims to compare the performance of the Support Vector Machine (SVM) multiclass classifier when entered with feature subsets generated from three different filter feature selection methods. The dataset is a clothing review text data taken from Kaggle. It contains multiple classes with 23486 review instances. Since filter selection methods utilize ranking, CHI, IG, and Relief were able to select and rank almost the same number of features. When the data was fed into SVM, CHI garnered 66.84% accuracy, IG got 32.90% accuracy while Relief obtained 29.69% accuracy. Experiments on SVM showed that among the generated data subset from the three filter selection methods, the subset using CHI yielded higher accuracy, precision, recall and F1 score compared to others.