Classifying Malicious Documents on the Basis of Plain-Text Features: Problem, Solution, and Experiences

Jiwon Hong, Dongho Jeong, Sang‐Wook Kim · Applied Sciences · 2022

Cyberattacks widely occur by using malicious documents. A malicious document is an electronic document containing malicious codes along with some plain-text data that is human-readable. In this paper, we propose a novel framework that takes advantage of such plaintext data to determine whether a given document is malicious. We extracted plaintext features from the corpus of electronic documents and utilized them to train a classification model for detecting malicious documents. Our extensive experimental results with different combinations of three well-known vectorization strategies and three popular classification methods on five types of electronic documents demonstrate that our framework provides high prediction accuracy in detecting malicious documents.

Read the paper · More papers on PaperTik