Large-scale exploration of feature sets and deep learning models to classify malicious applications
Tristan Vanderbruggen, John Cavazos · 2017
In recent years, researchers have shown that deep learning (DL) can be used to construct highly accurate models to solve many problems. However, training DL models requires large datasets and vast amounts of computation. With millions of malware variants being created every day, we contend that there is plenty of data to build deep learning models to classify malicious applications. However, finding the best DL model for this task requires exploring a wide range of methods to characterize malware and a variety of different DL models that can be used to classify malicious applications. To the best of our knowledge, no work has been presented that explores the large malware characterization space together with the variety of different DL models that could be brought to bear to this problem. In this paper, we present our work on exploring a large set of features and different DL models for the problem of malware family classification. To make this possible, we built a scalable machine-learning platform on Amazon Web Services (AWS). This platform made it possible to train many DL models concurrently on thousands of machines while collecting accurate data and performance information at regular checkpoints for reliability. We used this platform to evaluate two hundred DL models for eleven different malware characterizations. These characterizations include seven novel graph-based characterizations of the structure of executable code. While state-of-the-art malware characterizations yield 13.8% error-rate, our novel graph-based characterizations make less than 6.3% errors.