A Machine Learning Approach to Threat Hunting in Malicious PDF Files
Haydar Teymourlouei, Vareva E. Harris · 2023
This paper investigates the effectiveness of machine learning models for cyber threat hunting. Four distinct machine learning algorithms are used: Support vector machine (SVM), k-nearest neighbors (KNN), multi-layer perceptron (MLP), and random forest (RFC). We also analyze the behavioral patterns associated with cyber threats, which can achieve an insight for improving threat hunting systems. The dataset used in the work is specific to threats associated with malicious PDF files. The ML models were tasked with classifying benign and malicious PDFs. The behavior analysis investigates patterns related to the general and internal structure of the PDFs. In the results, we demonstrate the RFC classifier had an accuracy over 99% and loss below 0.05. The other classifiers had at least 95% accuracy and less than 0.2 loss. When using the full dataset, the MLP model had significantly larger computational overhead than the others. The method presented here has the potential to enable real-time threat hunting for a variety of cyber-security applications.