Detecting Phishing Websites using Naive Bayes Classification

Ahmed Rohan Talukder, Faria Alam, Sazia Tabasum Mim, Md. Abdul Al Emon · 2024

Phishing is one form of cyber attack used to obtain sensitive information from targeted individuals. Through this process, the attacker masquerades as a domain similar to an official legitimate website. Perpetrators can execute attacks with minimal expenses and effort, allowing them to initiate numerous assaults rapidly. The swift pace of phishing underscores the importance of automated detection processes in ensuring the protection of Internet users. Recently, many promising approaches, such as the content-based approach, heuristic-based approach, Fuzzy’s rule-based approach and machine-learning approach have been used to distinguish between phishing and legitimate ones. The existing studies do not focus on a probabilistic approach. Therefore, this study uses the Naive Bayes classification to detect phishing sites. The study compares three naive Bayes approaches: Bernoulli Naive Bayes, Multinomial Naive Bayes, and Complement Naive Bayes. The evaluation result of the Complement Naive Bayes Classifier and Multinomial Naive Bayes is superior to Bernoulli Naive Bayes. Both of the classifiers give 96 percent accuracy. As for detecting phishing websites, text classification and feature extraction are essential tasks; Naive Bayes classification provides an appropriate output.

Read the paper · More papers on PaperTik