Semantic Feature Discovery of Trojan Malware using Vector Space Kernels

John A. Musgrave, Carla Purdy, Anca Ralescu, David A Kapp, Temesgen Kebede · 2020

Malware analysis techniques without context or abstraction can be easily evaded by small changes in syntax. Semantic features represent an underlying pattern of abstraction in a program. Vector space models of malicious program instructions enable malware analysis via supervised document analysis and unsupervised cluster analysis. Document level features can be used for supervised learning, and malicious program semantics can be captured by both the eigenvalue decomposition of term frequencies and community discovery in the feature vector space. To obtain a fine-grained view of malicious program semantics intra-document opcode structures can be viewed without class labels for unsupervised clustering, as well as used in aggregate with class labels for malicious topic discovery via Singular Value Decomposition and to create a fine-grained semantic kernel. These patterns can be used to train a suitable classifier based on semantic features. The selection of these features can be evaluated by several criteria, including their embedded classifier performance. This paper presents preliminary results of classifiers trained to recognize static features of malicious documents and subgroup discovery via clustering algorithms to identify relevant features based on a semantic interpretation.

Read the paper · More papers on PaperTik