MalwareCLIP: Towards a Scalable and Explainable Image-Based Malware Classification

Yi Tan, Koh Ting Yew, Lee Joon Sern · 2023

Accurate and prompt malware classification allows cybersecurity defenders to react to threats precisely and quickly, ameliorating their impact. However, as attackers become more sophisticated, defenders face an escalating challenge. The existence of Malware-as-a-Service exacerbates the situation, making it easier to launch cyber-attacks. At present, malware classification has shown promising results through the use of deep learning. This is exemplified by converting malware binaries to their image equivalent, thereafter employing well-known Convolutional Neural Networks to perform classification. Though yielding high accuracy, they might not be well-suited in a highly dynamic domain, where novel malware samples are authored frequently. In this work, we attempt to bridge this gap by proposing a novel approach to train classifiers that are competitive to baselines while being scalable and explainable. We propose MalwareCLIP, gaining inspiration from prior art. Instead of using one-hot or integer-encoded labels for training, we use natural language training labels. We evaluated our approach on a real-world dataset with more than 20 malware families and show that MalwareCLIP is a promising direction for malware classification from our competitive classification results compared to baselines. Furthermore, our results show that MalwareCLIP can adapt to new classes without the need to change the architecture, thereby providing scalability. Finally, and more importantly, MalwareCLIP can also provide additional insights into family similarity, offering explainability.

Read the paper · More papers on PaperTik