Text Phishing Detection System using Random Forest Algorithm

Radhika Rajoju, V V R S S S Sathvika, Guthikonda Narayana Sai Smaran, Chettipelli Tejashwini, Gangasani Aditya Reddy · 2024

This research paper presents a comprehensive exploration of advanced text classification methodologies, specifically tailored for the intricate task of distinguishing phishing emails from legitimate ones. Employing cutting-edge natural language processing techniques, our study meticulously pre-processes the dataset, ensuring optimal refinement through meticulous steps such as null value elimination, text normalization, punctuation removal, and stop-word elimination. By harnessing sophisticated feature engineering techniques like CountVectorizer and TF-IDF Transformer, we seamlessly convert textual data into numerical representations, enabling our machine learning algorithms to operate with precision and efficiency. The experimentation phase rigorously evaluates a diverse range of classifiers, including Naive Bayes, Decision Trees, Logistic Regression, Random Forest, AdaBoost and KNN. Remarkably, our findings highlight Random Forest as the standout performer, boasting an impressive accuracy rate of 96%. This exceptional performance underscores the algorithm’s efficacy in accurately discerning between phishing attempts and genuine communications. As such, we advocate for the widespread adoption of Random Forest as the preferred model for email classification tasks, offering unparalleled accuracy and robustness essential for real-world deployment in cybersecurity contexts.

Read the paper · More papers on PaperTik