Enhanced phishing URL detection using hybrid feature-based machine learning method
J. Kumar · IET conference proceedings. · 2023
The rapid growth of the internet has led to increased usage of e-commerce, but it has also attracted cybercriminals who seek to steal personal information through various means, including phishing schemes. In these schemes, individuals are tricked into disclosing sensitive information through fake URLs. Due to the semantics-based attack strategy used in phishing, it can be difficult to distinguish between legitimate and phishing URLs, taking advantage of computer users' vulnerabilities. Although software companies offer anti-phishing systems that utilize blacklists, heuristics, visuals, and machine learning, they cannot prevent all phishing attempts. To address this issue, this research proposes five classification methods that use hybrid features, including natural language processing (NLP) and principal component analysis (PCA). The study finds that the Random Forest algorithm and XGboost utilizing NLP and word vector features outperform their competitors, achieving a 99.5% accuracy rate in classifying phishing URLs. These findings are significant as they demonstrate the potential of machine learning techniques, such as NLP and PCA, in improving the accuracy of anti-phishing systems. By incorporating these techniques into the classification process, it is possible to enhance the ability of such systems to differentiate between legitimate and phishing URLs. This can help in preventing cybercriminals from stealing sensitive information and provide a safer online experience for users.