PDF Malware Detection Using Machine Learning

Naragorn Tantikulvichit, Somkiat Kosolsombat, Chiabwoot Ratanavilisagul · 2024

In today's world, cyber-attacks are on the rise, and PDF files are commonly used as a means of attack. One common type of attack through PDF files is the covert embedding of dangerous malware and tricking users into clicking on malicious links. The victims may have their important data stolen and suffer various forms of damage. Detecting malware embedded in PDF files can help mitigate the harm to users. This project aims to study methods for detecting malware embedded in PDF files using machine learning techniques, with the best-performing model being Gradient Boosting, achieving an accuracy of 99% in 5-fold cross-validation and 97% using 10 features. Using Feature Selection Technique Voting from 3 Models consist of Logistic Regression, Random Forest, Lasso Regression, Each model will choose matching features to use for machine learning analysis. This project approach reduce feature for machine learning to detect PDFMalware.

Read the paper · More papers on PaperTik