Detecting Spam and Malware Using BERT and LLMs
Shahad Altamimi, Mohammed Mahmoud Ababneh · 2024
Spam and malware hinder the growth of IT and its use in many facets of contemporary life. This work employs directional Encoder Representations Transformers (BERT) and Large Language Models (LLMs) to enhance the detection of spam and malware. Researchers use cutting-edge Natural Language Processing (NLP) to detect malicious information with exceptional precision and resistance. BERT is an open-source NLP and Machine Learning (ML) framework. Text around unclear phrases help computers grasp context. It uses transformers, a Deep Learning (DL) model with dynamic weightings that connects all output and input elements. LLMs and BERT are used to identify SMS Spam and Malware using memory samples in this study. In the experiments, BERT demonstrated excellent performance on the datasets it used. By using BERT on SMS spam and malware datasets, accuracy of 100% and an F1-score of 99.99% for spam detection, alongside an accuracy of 100% and an F1-score of 100% for malware detection. Furthermore, the results demonstrate that BERT and LLMs have the potential to revolutionize cybersecurity, digital threat mitigation, and natural language processing classification.