Multimodal Phishing Detection on Social Networking Sites: A Systematic Review

Tandin Wangchuk, Tad Gonsalves · IEEE Access · 2025

Phishing is one of the most common cyberattacks, with the number of incidents increasing annually. Significant research interests have been generated in phishing emails, URLs, and websites over the past decade. Phishing attackers often target easy prey, and phishing on social networking sites (SNS) is increasing due to the popularity of easier communication using short text, images, and voice. Multimodal features can be used for phishing, yet there is relatively less research on phishing on SNS using multimodal features and models. This systematic review of the PRISMA and SEGRESS-guided literature aimed to uncover the techniques used for multimodal phishing detection.Acomprehensive literature search in Scopus, Web of Science, IEEE Xplore, and ACM Digital Library was conducted, including studies published in English from 2018 to 2025. A total of 20 studies are included in the study out of 74 records returned. Twelve articles were included from references of the selected studies. The study reveals that HTML content, URLs, and visual elements are used as multimodal features to classify phishing URLs and websites. These studies used deep learning models, including CNN, RNN, LSTM, and MLP, which produced promising results. However, challenges persist related to resource constraints, adversarial attacks, data quality and availability, false positives and negatives, and integration with existing SNS frameworks. These multimodal studies show promise, but require adaptation to SNS-based phishing attacks.

Read the paper · More papers on PaperTik