Unmasking the Phishermen: Phishing Domain Detection with Machine Learning and Multi-Source Intelligence

Radek Hranický, Adam Horák, Jan Polišenský, Kamil Jeřábek, Ondřej Ryšavý · 2024

In the digital landscape, phishing attacks have rapidly evolved into a major cybersecurity challenge, posing significant risks to individuals and organizations. This short paper presents our preliminary research on detecting phishing domains. Our approach amalgamates intelligence from multiple sources: DNS servers, WHOIS/RDAP, TLS certificates, and GeoIP data. We created a rich 15.8 GB dataset of information about benign and phishing domains, from which we derived a comprehensive 80-feature vector for training and testing machine learning classifiers. We propose preliminary results with a fine-tuned XGBoost model, achieving 0.9716 precision rate, 0.9540 F-1 score, and false positive rate of 0.23%.

Read the paper · More papers on PaperTik