Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails
Wei Hao, Van Tran, Vincent Rideout, Zixi Wang, AnMei Dasbach-Prisk, M. H. Afifi, Junfeng Yang, Ethan Katz-Bassett, Grant Ho, Asaf Cidon · 2025
The rapid adoption of large language models (LLMs) has fueled speculation that cybercriminals may utilize LLMs to improve and automate their attacks.However, so far, the security community has had only anecdotal evidence of attackers using LLMs, lacking large-scale data on the extent of real-world malicious LLM usage.In this joint work between academic researchers and Barracuda Networks, we present the first large-scale study measuring AIgenerated attacks in-the-wild.In particular, we focus on the use of LLMs by attackers to craft the text of malicious emails by analyzing a corpus of hundreds of thousands of real-world malicious emails detected by Barracuda.The key challenge in this analysis is determining ground truth: we cannot know for certain whether an email is LLM or human-generated.To overcome this challenge, we observe that, prior to the launch of ChatGPT, email text was almost certainly not LLM-generated.Armed with this insight, we run three state-of-the-art LLM detection methods on our corpus and calibrate them against pre-ChatGPT emails, as well as against a diverse set of LLM-generated emails we create ourselves.Since the launch of ChatGPT, all three detection methods indicate that attackers have steadily increased their use of LLMs to * Work done at Columbia University.