Malicious Word Document Detection Based on Multi-View Features Learning
Xiaofeng Lu, Fei Wang, Zifeng Shu · 2019
As users become more cautious about executable programs, malicious document attacks have become popular, and the proportion of malicious document attacks has become higher and higher. Therefore, detecting malicious documents is one of the most important tasks in information security. Microsoft's Word Document has become one of the main ways for attackers to use. Attackers often insert VBA programs into a Word or use CVE to launch attacks. The more dangerous attack is to exploit the vulnerability of the document reader software to enable Shellcode to execute. The exploitation of vulnerability is related to the content and structure of the document. Based on the above reasons, we present a method to analyze Word documents from four independent views: VBA functional words, Ole file object formats, structure paths, and specification errors. This paper studies the features and feature extraction methods of malicious documents in these four views. Based on the features in the above four views, decision trees and random forest machine learning algorithms are used to study and classify the documents. Our classification method detected malicious Open-Xml-Document with a high average recall rate (97.38%) and a low FPR rate (0.7%). In the detection of malicious MS-DOC documents, the average recall rate is 97.9% and FPR rate is 1.4%.