Hybrid Filter-Wrapper Text Feature Selection Technique for Text Classification
Osamah Mohammed Alyasiri, Yu–N Cheah, Ammar Kamal Abasi · 2021
Feature selection is one of the most common and vital data preprocessing techniques in text classification. It is used to reduce the high-dimensional set of features in any dataset by eliminating irrelevant, redundant, and noisy features that are unnecessary from the classification point of view, since it makes it difficult for classifiers or machine learning techniques to produce correct classification results. This research work has accomplished a twofold objective to achieve these goals. Firstly, it aimed to select top-N features with the highest-ranking features using the Information Gain (IG) filter approach. Secondly, it intended to reduce the text feature set obtained by IG using the gray wolf optimizer (GWO) search strategies in the wrapper approach to finding the most informative subset of text features. The Naive Bayes (NB) classifier was introduced to evaluate the performance of this proposed method. Experimentations were carried out on nine benchmark document datasets. Based on the evaluation measures, the results show that the IG-GWO could be used as an alternative method to address the text FS problem compared to the state-of-the-art feature selection algorithms.