Feature Selection and Classification Performance: A Multi-Dataset Comparative Analysis Using Boruta Algorithm and Random Forest

Ikhlass Boukrouh, Faouzi Tayalati, Abdellah Azmani · Proceedings of the World Congress on Electrical Engineering and Computer Systems and Science · 2024

Dimensionality reduction is crucial for managing high-dimensional datasets in machine learning, reducing complexity and overfitting.This study evaluates the efficiency of classification models without and with feature selection using the Boruta algorithm with Random Forest classifiers across three distinct datasets.Feature selection aims to improve model accuracy and interpretability by retaining only the most significant features.The three datasets were evaluated using full and reduced feature sets by comparing accuracy, precision, recall, and F1-score.Results show that feature selection significantly enhances model performance.For Dataset 1, accuracy improved by 1.06%, precision by 3.23%, recall by 3.46%, and F1-score by 3.36%.Dataset 2 saw increases in accuracy by 0.46%, precision by 2.36%, recall by 4.82%, and F1-score by 5.42%.Dataset 3 showed no significant changes, with both configurations yielding similar performance metrics.These findings confirm that the Boruta algorithm effectively enhances classification performance by reducing dataset dimensionality and retaining key features, especially in datasets with irrelevant features.However, when all features are relevant, the benefits of feature selection may be minimal.

Read the paper · More papers on PaperTik