Towards Explainability Using ML And Deep Learning Models For Malware Threat Detection

Mattaparti Satya Chandana Snehal, Veeraboina Nagoor, S. Rohit, Sneha Raghunandan, Senthil Kumar Thangavel, Kartik A. Srinivasan, Pratyul Kapoor · 2024

In response to the escalating threat of Android malware, this research proposes a hybrid model for malware detection and classification using a combination of machine learning (ML) and deep learning techniques. Leveraging Random Forest for feature extraction, the model efficiently processes the complex feature spaces present in the CICMaldroid dataset. Simultaneously, the Long Short-Term Memory (LSTM) CNN, Bidirectional LSTM (BILSTM), Bidirectional LSTM (BILSTM) CNN are employed for further analysis. The proposed approach integrates predictions from both the machine learning and deep learning models to achieve comprehensive results. Data preparation involves outlier removal and Synthetic Minority Over-sampling Technique (SMOTE) for class imbalance. Random Forest analysis guides feature selection, optimizing the model’s efficiency. Cross-validation ensures robustness, and Explainable AI techniques, specifically SHAP, enhance interpretability. Experimental results demonstrate the model’s effectiveness, achieving high overall accuracy and specific identification of Android malware types. The proposed methodology provides a streamlined and transparent approach to early malware detection, contributing to a deeper understanding of the decision-making process behind each prediction.

Read the paper · More papers on PaperTik