Self-Learning Multimodal AI Framework for Toxic Content Moderation in Real-Time Retail Platforms
Arjun Sirangi · 2025
As real-time retail platforms change to include various modes of shopping, managing toxic content for example, hate speech and misinformation has become more complicated. Conventional content moderation methods which use only a single method and are based on rules, find it difficult to handle the mixed and complex nature of today’s user-generated media. To moderate images, text and videos, the paper presents a Self-Learning Multimodal AI Framework that benefits from Large Language Models, computer vision and audio analysis. The model combines Convolutional Neural Networks (CNNs), Transformers and Bi-directional Attention Mechanisms (BAM) which help in understanding both the meaning and appearance of items. We make sure the system is adaptive, so it can achieve more accurate moderation as time goes by, lessening the amount of manual supervision needed. It also makes choices easy to explain and uses user feedback to ensure the system is fair and justified. Experiments on many kinds of data show that the framework outperforms existing systems in terms of precision, recall and response speed. With this research, it is possible to design scalable and fair ways to monitor AI content on large online platforms. It gives retailers the means to be more vigilant in dealing with toxicity as it happens online.