Attack Behavior Extraction Based on Heterogeneous Threat Intelligence Graphs and Data Augmentation

Jingwen Li, Ru Zhang, Jianyi Liu · 2024

Recently, cyber attacks have become increasingly complex and diverse. Threat intelligence is being systematically integrated to enhance threat detection capabilities. Tactics, Techniques, and Procedures (TTPs), as an advanced form of threat intelligence, provide detailed information about attackers’ methods of operation. This information holds significant value for understanding attacker behavior patterns and implementing proactive defense measures. In recent times, machine learning methods have been applied to identify TTPs. However, annotating TTPs requires experts with profound knowledge of offensive and defensive tactics. This results in a limited amount of available labeled data, limiting the effectiveness of machine learning models. To address this issue, we propose a method for extracting attack behavior based on heterogeneous threat intelligence graph and data augmentation. Firstly, a large amount of unlabeled threat intelligence text is utilized to train a masked language model based on Bidirectional Encoder Representation from Transformers (BERT). By predicting masked tokens within the annotated text, new text is generated that aligns with the textual characteristics of threat intelligence. Secondly, to effectively leverage contextual information in threat intelligence, we construct a Heterogeneous Threat Intelligence Graph (HTIG), modeling threat entities and their associated relationships. Simultaneously, to alleviate the issue of feature sparsity, a graph attention network is employed to learn embedded representations of nodes in the HTIG, facilitating the extraction of TTPs. Experimental results on the ATT&CK dataset indicate that our approach significantly enhances the efficiency of TTP identification.

Read the paper · More papers on PaperTik