Improvement of Machine Learning Algorithms with Hyperparameter Tuning on Various Datasets
Akmar Efendi, Iskandar Fitri, Gunadi Widi Nurcahyo · 2024
The rapid growth of data in the digital era has made classification techniques a critical component of machine learning, particularly in supervised learning methods. These techniques enable computers to learn from labeled data and make predictions on unseen data by identifying patterns in the training data. However, despite the reliability of algorithms like Support Vector Machine (SVM) and Naïve Bayes, their performance can be suboptimal when dealing with complex datasets. This study explores the effectiveness of various machine learning models enhanced by optimization techniques across different types of datasets. Specifically, it evaluates the impact of using SMOTE for data balancing and the integration of hyperparameter tuning with Optuna and XGBoost to improve the performance of SVM and Gaussian Naive Bayes (GNB) classifiers. The experiments were conducted on three datasets: academic data from Universitas Islam Riau, sentiment analysis data from Twitter, and medical data from a diabetes dataset on Kaggle. The results demonstrate that the combined use of SMOTE, SVM, and Optuna can achieve perfect accuracy on academic data, while similar improvements were observed in Twitter and diabetes data when GNB was combined with XGBoost and Optuna. These findings underscore the potential of optimization techniques in enhancing machine learning model performance across diverse applications, paving the way for more sophisticated and effective classification models in the future.