Comparison of Machine Learning Methods for Android Malicious Software Classification based on System Call
Mochammad Isa Anshori, Farhanna Mar’i, Fitra Abdurrachman Bachtiar · 2019
The development of the Android operating system is very rapid accompanied by the development of various types of malicious software (malware). The malware application can enter automatically into an Android device in an unintentional way by Android smartphone users so that there are many cases of data theft that are very detrimental to the user. In this study, malware detection will be based on the system call feature on Android using several machine learning methods, namely Support Vector Machine (SVM), Naïve Bayes, Decision Tree, Random Forest, Log Regression, and K-nearest Neighbor (KNN). The purpose of this study is to find out the machine learning method that can provide the best value of accuracy, TPR, and FPR in resolving the problem of malware detection on android by classification of types of malware using a system call on Android. Based on the results of this study, it can be seen that the Random Forest (RF) method can classify malware in an android system by conducting early detection that produces an accuracy value of 76%, Random Forest has proven to have reliable performance in case of classification and also has advantages such as fast computation time and high accuracy also proved to be better than other machine learning methods, namely SVM, Naïve Bayes, Decision Tree, Log Regression, and K-nearest Neighbor (KNN), which each produced an accuracy value of 71.67%, 66.83%, 69.33%, 70.83% and 71.67%.