Spear Phishing Detection

Shibayan Mondal, Samrajnee Ghosh, Achiket Kumar, SK Hafizul Islam, Rajdeep Mohan Chatterjee · 2022

In the world today, security in the digital realm is of utmost importance. The path to achieving security begins with detecting the kind of attack and identifying patterns in the staged ways. These attacks are often termed as “phishing.” It can be described as a social engineering attack involving sophisticated attack vectors that can be used to steal sensitive information from a victim, to be more precise as per the definition of phishing. In phishing, the attacker disguises himself as a trustworthy entity, uses several tactics to win the victim’s trust, and then convinces the victim to reveal the requisite information to commit the fraud. This kind of attack can be done by sending malicious emails containing malware that can leak information when installed in the victim’s system. Such attacks can further be spread through the hacked systems to other victims. Phishing can again be classified into several types: Spear phishing, email phishing, whaling, smishing, vishing, and angler phishing. Among these, this chapter primarily concentrates on spear phishing. Spear phishing is a kind of email or social media scam that generally targets some specific individual or institution. Most of the time, the target happens to be influential. In spear phishing, the attacker gathers as much information as possible about the target to tailor the attack for that particular victim. However, many times, these attacks are found to exhibit a similar pattern. So, detecting that pattern in any email or social media post can predict any possible threat in the same. This chapter mainly focuses on finding these patterns and labeling posts as possible spear-phishing attacks using machine-learning (ML) algorithms. The classification algorithms categorize structured or unstructured data into a given number of decision classes. There are different classification algorithms, like decision tree, support vector machine, multinomial Naive Bayes, logistic regression, and K-nearest neighbors. Besides, to improve the accuracy of the classification of the data, ensemble methods are used. In machine learning, several techniques are combined to form ensemble methods to balance variance and bias or improve predictions. These ensemble methods can further be divided into sequential ensemble methods like AdaBoost, Stacking, and so on and parallel ensemble methods like Random Forest. Here, both ways have been used to have a comparative study between the accuracies of classical machine learning algorithms and the ensemble techniques. Furthermore, it is observed from our study that the ensemble approaches outperform the traditional machine learning techniques in predicting spear phishing in an email.

Read the paper · More papers on PaperTik