Automated Threat Detection in the Dark Web: A Multi-Model NLP Approach
Abhay Kamath, Aditya Joshi, Aditya Sharma, Nikhil R Shetty, Lenish Pramiee · 2025
The dark web, recognized for its anonymity as well as for its encryption, poses a number of significant challenges for law enforcement and cybersecurity due to its illegal activities, which may include cybercrime and illegal trade. Classic monitoring techniques simply cannot cope with the dark web's unorganized, profusely scaled data. This paper intends to present an automated threat detection system based on transformer architecture models—BERT, DistilBERT, and DarkBERT—to maximize classification performance and confront threats in real-time. These models are fine-tuned on diverse datasets, DNRTI, Agora, and CoDA respectively, to classify and analyze threats effectively. Our approach integrates machine learning with a cloud-based infrastructure, namely Firebase, that can scale up to perform efficient classification, visualization, and threat reporting. This system proves superior to traditional deep learning models by achieving accuracy above 95 % on several threat classification tasks. In contrast to previous research works that stress the use of one-specific NLP model, our system synergically integrates multiple NLP models in layers to provide an enhanced classifier for better detection accuracy. This framework brings together automation for data ingestion, classification, and visualization for a realistic and scalable solution for cybersecurity practitioners keeping an eye on threats on the dark web.