Towards an effective and efficient malware detection system

C.-T.D. Lo, Pablo F. Ordóñez, Cepeda Mora Carlos · 2016

The ubiquitous advance of technology used on the Internet, computers, smart phones and tablets has been conducive to the creation and proliferation of cyber threats resulting in attacks that have grown exponentially. Consequently, anti-virus companies and researchers have developed new approaches for dealing with discovering and classifying malware. Among these, machine learning and big data technologies have been used for feature extraction, detection, and clustering of cyber threats. In this paper a dataset of malware and clean files (goodware) was created and analyzed from the static and dynamic features provided by the online framework VirusTotal. The purpose is to select the smallest number of features that keep classification accuracy as high as possible in order to decrease the use of resources for monitoring as well as extracting features and the time for detection. In this research, it was found that “9” features are enough to distinguish malware from “goodware” files with an accuracy of 99.60%. Selecting the most representative features for malware detection relies on the possibility of creating an embedded program that monitors the processes executed by the operating system (OS) and looks for the characteristics that match malware behavior. In addition, classification algorithms such as Random Forest (RF), Support Vector Machine (SVM) and Neural Networks (NN) were used in a novel combination that not only showed an increase in accuracy, but also in the training speed from hours to just minutes. Finally, the trained model (which was trained with a dataset of malware samples seen before September 2015) was tested on a new dataset of malware samples seen by first time between October 2015 and June 2016 and showed that the model is still effective for detection of unseen malware files.

Read the paper · More papers on PaperTik