Enhancing Malware Detection Accuracy: A Comparative Analysis of Machine Learning Models with Explainable AI

Manoj Kumar, S Darshan, Sai Harshavardhan, Ashwini Kodipalli, Trupthi Rao · 2024

The challenge of detecting malware in cybersecurity necessitates the exploration of diverse machine learning models for effective solutions. In this investigation, we assess the performance of several models, including Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Logistic Regression, Decision Trees, Random Forests, Randomized Search CV, Grid Search CV, CatBoost, ADA Boost, XGBoost, and Gradient Boosting, using a comprehensive dataset. Our study reveals that CatBoost emerges as the most effective model, achieving an exceptional 99% accuracy in malware prediction. Furthermore, we utilize Lime and Shap explainability techniques to delve into the decision-making process of CatBoost, shedding light on its predictive mechanisms and enhancing transparency in malware detection. By emphasizing the importance of diverse machine learning models and the significance of explainable AI techniques in cybersecurity, our research contributes to the advancement of malware detection methodologies. It underscores the critical role of transparent and interpretable AI systems in fostering trustworthiness and effectiveness in cybersecurity applications.

Read the paper · More papers on PaperTik