CrossMemeNet: A Cross-Modal Attention Framework for Meme Sentiment Analysis Using CLIP and BERT

V Akshaya, Sivanantham S, Vijaya Krishnamoorthy, R. Senthil Kumaran, Renuka Devi S · 2025

Internet memes have become a new means of communication that include a blend of images and texts so as to express sentiment and emotions. Standard approaches to perform sentiment analysis are no longer efficient when the mems are multiform, which necessitates more complex possibilities for identifying them. To this end, CrossMemeNet is introduced in this paper as a new cross-modal framework grounded on the merits of CLIP and BERT for meme sentiment analysis. The integration is based on a two-streasmuner approach, in which CLIP processes the visual content with Vision Transformer and BERT the textual content with contextual embedding. All these features are connected by crossmodal attention that adaptively adjusts the contribution of all the modalities. The adopted framework is tested on the Memotion 7K corpus following a strict training regimen accompanied by extensive regularization. Experimental results pointed out that the proposed CrossMemeNet has achieved an accuracy measure of 85.27% and loss of 1.1234 in sentiment classification. The model is shown to have very high predictive confidence in all the sentiment categories with confidence percentages within the range of 87.30% – 91.65% on test samples. Training dynamics also demonstrate constant enhancement of both parameters; training accuracy ranges between 65% and 95% and validation accuracy is at 85 %. Based on these results, we can conclude that this subtle interplay between visual and textual inputs can be controlled by a careful architectural design of the advanced vision-language models. A study on the result shows that it has the potential to be used in its wider sense in the analysis of the content of the social media and the automatic identification of the sentiment analysis.

Read the paper · More papers on PaperTik