Feature Selection and Improving Classification Performance for Malware Detection

Carlos M. Cepeda, Dan Lo Chia Tien, Pablo F. Ordóñez · 2016

After analyzing the advance of technology, it is clear that use of the Internet, computers, smart phones and tablets has become ubiquitous and therefore, the creation and proliferation of cyber threats and attacks has grown exponentially. Consequently, Anti-Virus companies and researchers have developed new approaches for dealing with discovering and classifying malware. Among these, machine learning and Big Data technologies have been used for feature extraction, detection, and clustering of cyber threats. In this paper, we created and analyzed a dataset of malware and clean files (goodware) from the static and dynamic features provided by the online framework VirusTotal. The purpose is to select the smallest number of features that keep the classification accuracy as high as possible given that the training execution time increase in polynomial time with respect to the number of features. In this research, we found that "9" features are enough to distinguish malware from "goodware" files within an accuracy of 99.60%. Selecting the most representative features for malware detection relies on the possibility of creating an embedded program that monitors the processes executed by the OS looking for the characteristics that match malware behavior. Thus, feature selection was made taking the most important features that keep the accuracy high and allows the creation of monitoring malware detection programs with a low overhead cost. In addition, classification algorithms such as Random Forest (RF), Support Vector Machine (SVM) and Neural Networks (NN) were used in a novel combination that not only showed an increase in accuracy, but also in the training speed from hours to just minutes.

Read the paper · More papers on PaperTik