Feature Selection Implementation on High-Dimensional Data using Firefly Algorithm

Tessy Badriyah, Iwan Syarif, Danang Djehan Lukmanto Prakoso · 2023

High-dimensional data often presents a challenge in machine learning tasks due to the curse of dimensionality. Feature selection is a common technique used to overcome this problem by reducing the number of features and selecting the most relevant ones. In this research, we propose the use of the firefly algorithm for feature selection in high-dimensional data. The aim of this study is to evaluate the performance of the firefly algorithm in selecting informative features and compare it to other feature selection algorithms. We used several datasets with different numbers of features and instances to test the performance of the firefly algorithm. The results demonstrated the capability of the firefly algorithm to enhance accuracy after feature selection process: in the Breast Cancer dataset, it improved from 0.959 to 0.969; in the Lymphoma dataset, it improved from 0.479 to 0.575; in the arrhythmia dataset, it improved from 0.622 to 0.714; and in the leukemia dataset, it improved from 0.621 to 0.733. However, for the Central Nervous System (CNS) dataset, the accuracy remained the same before and after at 0.633. We also compared the results of the firefly algorithm with three other well-known feature selection algorithms: the genetic algorithm, particle swarm optimization, and ant-bee colony algorithm. Using performance metrics such as accuracy, precision, and recall, the firefly algorithm proved to be efficient in handling high-dimensional datasets with satisfactory results in several cases. It outperformed the other algorithms, particularly the ant-bee colony algorithm, which also surpassed the others in certain cases.

Read the paper · More papers on PaperTik