Evaluating Performance Maintenance and Deterioration Over Time of Machine Learning-based Malware Detection Models on the EMBER PE Dataset
Colin Galen, Robert JC Steele · 2020
Academic studies evaluating the performance of machine learning-based malware detection models have in numerous cases demonstrated high accuracy measures, based upon training and evaluation on large datasets of malware/goodware examples. However, an important consideration is how well those high performing models will maintain their performance into the future in real-world settings, after the time of training and in the face of the varying malwares and goodwares subsequently encountered over time. In this work, we consider model performance maintenance on a large time-ordered dataset of the features extracted from one million malware/goodware samples. We train and evaluate a large range of models, and see a wide variation in model performance maintenance for different base models, even where initial performance of models is similar. We investigate in further detail those classes of model and specific models that show highest performance maintenance and discuss the significance of these findings for building robust real-world machine learning-based malware detection systems. This is the first work we are aware of to evaluate the relative rate of performance deterioration of different machine learning-based malware detection models evaluated utilizing such a large real-world malware/goodware dataset.