Phishing URL Detection Using Deep Learning with CNN Models
Alsadig Hadi Alsadig, Md Oqail Ahmad · 2024
This research study aims to enhance the performance of phishing URL detection models through a series of strategic improvements. Addressing common challenges such as dataset imbalance, data variability, and the need to minimize false positives and negatives, this study implements several key methodologies. To counteract dataset imbalance, we utilize SMOTE (Synthetic Minority Oversampling Technique) to balance the dataset, which improves the model's capacity to accurately identify phishing URLs. This study applies feature scaling and selection techniques, with SelectKBest identifying the top 16 features based on their F-value scores, to reduce data dimensionality and focus on the most relevant attributes. This study trains two models: A Convolutional Neural Network (CNN) for its efficacy in text classification and a Random Forest Classifier (RF) for its robustness against class imbalance and ability to handle complex data patterns. Hyperparameter optimization is performed using GridSearchCV, and the decision threshold is adjusted to 0.4 to potentially lower false positives by refining the conversion of predicted probabilities to binary predictions. The experimental results show that both models achieve high accuracy, with the Random Forest model reaching 99.79% and the CNN model achieving 99.75%. These findings highlight the effectiveness of the RF and CNN models in enhancing online security and mitigating phishing attacks. The implemented strategies significantly improve the model's accuracy and reliability in detecting phishing URLs, ensuring reduced rates of false positives and negatives.