Hierarchical multimodal robust spam detection using Large Language Models and Convolutional Networks

Rania Mkhinini Gahar, Adel Hidri, Olfa Arfaoui, Minyar Sassi Hidri · Procedia Computer Science · 2025

With the growth in the use of email and social media, spam has become a major challenge. With the rise of multimedia technologies, the prevalence of multimodal spam containing a mixture of text and images has significantly increased. However, most of the methods proposed to detect spam in the past are mainly based on text analysis. The development of a multimodal approach to spam filtering is therefore of paramount importance. The paper aims to develop an improved method for multimedia spam detection using hierarchical and multimodal message analysis, combined with deep learning (DL). Our approach is based on extracting several representative characteristics from multimodal data (text, links, images) using models based on large language models and convolutional networks. This aims to obtain a fine-grained semantic representation with focus on the key elements of the messages for more effective classification of multimedia spam. We experimented our method on a large corpus and used qualitative and quantitative analysis to compare accuracy and robustness. The experiments demonstrate the model’s capability to detect effectively complex spam strategies, reflecting significant improvement in spam identifying techniques.

Read the paper · More papers on PaperTik