Victim or Attacker? A Multi-dataset Domain Classification of Phishing Attacks

Sophie Le Page, Guy-Vincent Jourdan · 2019

In “phishing attacks”, phishing websites disguised as trustworthy websites attempt to steal sensitive information from end users. Remediation options differ depending on whether the phishing website is hosted on a legitimate but compromised domain, in which case the domain owner is also a victim, or whether the domain itself is maliciously registered by an attacker. We propose here a novel machine-learning domain classifier, introducing features based on the internet presence and history of a domain, using only publicly available information. Using a phish feed and malicious domain feed from the Anti-Phishing Working Group (APWG), evaluation of our domain classifier achieves 94% accuracy on future malicious domains, while maintaining 88% and 92% accuracy on malicious and compromised datasets respectively from two other sources. To increase our training set we introduce a semi-supervised technique to label part of APWG's phish feed. For the rest of the feed we use our classifier and show that 62% of the websites hosting attacks are compromised while the remaining 38% belong to the attackers. The result of this research is a tool which gives some important insights on phisher's current modus operandi. It also provides a quick mechanism, entirely based on freely available data, to assess crucial information about the server and better respond to an attack by having it taken down as quickly as possible.

Read the paper · More papers on PaperTik