Identification of Malicious PDFs Using Convolutional Neural Networks

Umm-e-Hani Tayyab, Muhammad Zain, Faiza Babar Khan, Muhammad Hanif Durad · VFAST Transactions on Software Engineering · 2022

With multiple possible carriers of malware, one of the most targeted file formats by malware writers is PDF format due to its inherent shortcomings as well as the limitations of PDF readers. The PDF-based attack is one of the most common attacks among document-based attacks. Some of the most luring features of PDF files include their widespread use which has replaced Word documents quite dominantly, secondly the ease of crafting a malicious PDF, and above all its capability of containing javascript. Initially, researchers focused on identifying malicious PDFs by comparing the structural changes between benign and malicious PDFs. Malware writers started using the evasion techniques to hide the malware so that structural changes can be hidden. Researchers started exploiting the benefits of Artificial Intelligence to combat these evasion techniques. We developed a convolutional neural network, trained to classify PDFs as benign and malicious using the byte level information, on a dataset of approximately 21,000 files. Our model can detect malicious PDFs without the human intervention and manual feature extraction. Moreover, our model detects malicious PDFs with 97% accuracy.

Read the paper · More papers on PaperTik