Criminal Email Detection Using Innovative Large Language Models and Data Augmentation
Levi Pittman, Mohamed I. Ibrahem, Mahmoud S. Abouyoussef · 2025
The increasing use of email as a communication tool has provided new opportunities for criminals to coordinate illegal activities. In response, law enforcement agencies have turned to advanced technologies to analyze email content and identify potential criminal behavior. Traditional methods for criminal detection often rely on known offenders, with some approaches using feature extraction and natural language processing (NLP) techniques to examine communication. However, with the recent advancements in large language models (LLMs), these tools have become significantly more powerful and versatile. This paper explores the use of LLMs to detect criminal content in emails. We demonstrate that LLMs offer improved accuracy in identifying emails that contain criminal information, thereby enhancing user protection. To evaluate the effectiveness of LLMs, we tested several models to classify criminal emails and addressed the challenge of imbalanced datasets using multiple data augmentation techniques. The results are compared to existing literature, illustrating the impact of LLMs and data augmentation in improving the classification of criminal emails.