Impact of Feature Selection and Interactions on Accuracy of Classical Machine Learning Algorithms: A Driver Behavior Case Study
Rabab Gamal, Mirvat Al-Qutt, Heba Khaled, Mahmoud Fayez, Said Ghoniemy · 2024
This study aims to investigate how the number of features and their interrelationships impact the accuracy of classical machine learning algorithms, focusing on driver behavior as a case study. As sensor data becomes increasingly available in modern vehicles, there is a growing interest in understanding driver behavior patterns for various applications, including safety and personalized services. However, the specific influence of different features and their interactions on the accuracy of machine learning models in predicting driver behavior remains unclear. To address this gap, the study utilizes a comprehensive dataset called UAH-DriveSet. Multiple subsets of the dataset are created by gradually adding or removing features, allowing for an examination of the impact of feature quantity. Additionally, feature engineering techniques are applied to explore the relationships and interactions between the features. The study employs a set of classical machine learning algorithms, including logistic regression, decision trees, random forests, support vector machines, naive Bayes, k-nearest neighbors, linear discriminant analysis, and AdaBoost, to model driver behavior based on various feature subsets. The performance of each algorithm is evaluated using accuracy and F1-score as the evaluation metric. The findings of the study reveal that the number and selection of features that were extracted after the preprocessing of the data set, play a crucial role in determining the accuracy of machine learning models for predicting driver behavior. The preprocessing is carried out using an autoencoder model for dimension reduction, and Singular value decomposition to get the most effective column to be converted to the alphabet using k-means, then extracting the motifs representing the new feature.