Enhancing Phishing Detection with AI: A Novel Dataset and Comprehensive Analysis Using Machine Learning and Large Language Models

Robin Chataut, Yusuf Usman, C. M. A. Rahman, Sohan Gyawali, Prashnna Gyawali · 2024

Phishing emails are a significant threat to organizations, with over $90 \%$ of cyber attacks starting from a malicious email. Despite built-in security measures, relying solely on these defenses can leave organizations vulnerable to cybercriminals who exploit human nature and the lack of tight security. Phishing emails, designed to deceive recipients into disclosing personal and financial information, represent a significant cybersecurity challenge. This paper introduces a comprehensive dataset curated explicitly for detecting phishing emails, featuring a collection of authentic and phishing emails. The dataset includes a broad spectrum of phishing techniques, such as sophisticated social engineering tactics, impersonation of reputable entities, and using urgent or threatening language to manipulate recipients. Phishing emails were collected to cover various scenarios, including financial fraud, account verification, and malware dissemination attempts. Our analysis involves a range of classical machine learning models alongside exploratory analysis with LLMs. The performance of these models was rigorously evaluated to furnish a comparative analysis of their detection capabilities. The dataset, one of the largest of its kind, offers a significant resource for researchers and cybersecurity professionals aiming to advance phishing detection methods. The dataset used in this research is publicly available, enabling further exploration and replication of the findings by the research community [1].

Read the paper · More papers on PaperTik