PDFalse: Evasive Malicious PDF Machine Learning Classifier
Agung Ngurah Gde K.T.D. Gusti, Girinoto · 2023
In recent years, detecting malicious PDF has become increasingly challenging. It is crucial to continuously improve mitigation and identification measures against malicious scripts embedded in PDF, particularly evasive malware. Evasive malware is designed to resemble harmless applications, making it difficult to detect. Despite this, only two publications in the past five years have addressed the issue of classifying evasive malicious PDF. This research seeks to fill this gap by determining the best classifier model for evasive malicious PDF. The approach involves benchmarking various machine learning models, including shallow and deep learning algorithms like Gradient Boosting, XGBoost, MLP, and Neural Networks. Each model’s performance is assessed based on the F1-score metric. Through a comprehensive process of data preprocessing, and model training with the CIC-Evasive- PDFMa12022 dataset, the Gradient Boosting algorithm emerges as the best-performing model with an impressive F1-score of 99.66%. The top-performing model is subsequently implemented into a.NET-based desktop application, named PDFalse, designed to classify PDF as safe or malicious in real-time, efficiently, and accurately. The effectiveness of PDFalse is further validated through black-box functional testing, reinforcing its practical utility in real-world scenarios.