Latent space refinement for unsupervised cyber threat text classification
Yue Wang, Richi Nayak, Md Abul Bashar, Mahinthan Chandramohan · Knowledge-Based Systems · 2025
• This paper proposes Latent Space Refinement (LSR), a novel unsupervised classification framework that integrates metric learning with clustering-based representation refinement, addressing the critical challenge of label scarcity in cyber threat intelligence (CTI). • LSR introduces a posterior regularisation strategy that aligns latent representations from Pretrained Language Models (PLMs) with an auxiliary TF-IDF-based distribution. This guides unsupervised adaptation to the target domain without any PLM fine-tuning, ensuring scalability and efficiency. • Extensive experiments on three CTI benchmarks demonstrate that LSR consistently outperforms state-of-the-art unsupervised and few-shot baselines in Accuracy and F1 score. • By enabling lightweight unsupervised domain adaptation, LSR offers a plug-and-play solution applicable to CTI, including other resource-constrained domains such as health and legal text classification. Text classification plays a critical role in Cyber Threat Intelligence (CTI) applications, where open-source text data is mined to identify patterns such as Indicators of Compromise (IoC), Tactics, Techniques and Procedures (TTPs), Named Entities and more. However, the dynamic nature of CTI makes traditional supervised machine learning classifiers impractical due to their reliance on large number of labelled training datasets. To address this, we propose Latent Space Refinement (LSR), an unsupervised method designed for CTI text classification. LSR introduces a posterior regularisation strategy where an auxiliary distribution derived from a TF-IDF feature space serves as signals to refine latent representations derrived from Pretrained Language Models (PLMs). By iteratively refining this latent space with clustering signals, LSR enables efficient similarity-based classification using only a few user-provided seed keywords. Extensive experiments on diverse CTI tasks, including both binary and multi-class classification, demonstrate that LSR consistently outperforms state-of-the-art unsupervised and zero-shot/few-shot methods in Accuracy and Weighted F1 score, all without tuning internal PLM parameters. This makes LSR a lightweight and PLM-agnostic solution for real-world CTI applications.