CLARITY: A Lightweight Multimodal Transformer for Harmful Content Detection

Gautam Siddharth Kashyap, Niharika Jain, Ebad Shabbir, Harsh Joshi, Usman Naseem, Jiechao Gao · IEEE Transactions on Artificial Intelligence · 2025

Social media platforms are vital to modern communication, but they also enable the spread of harmful content, such as hate speech and misinformation. Current detection models, while accurate, are often resource-intensive and unsuitable for real-time or resource-constrained environments. Moreover, even models that incorporate multilingual capabilities often fail to generalize effectively across different languages. To address this challenge, we propose CLARITY, a novel lightweight cross-modal transformer architecture designed for efficient and scalable harmful content detection. Unlike traditional models, CLARITY achieves faster processing while maintaining accuracy, making it accessible to a wider range of platforms and devices. CLARITY integrates text, image, and audio modalities to capture complex, multimodal interactions that enhance detection across diverse content types. By employing contrastive learning, CLARITY accurately distinguishes between reclaimed language and genuinely harmful content, significantly reducing false positives and promoting inclusivity, particularly for marginalized communities. Additionally, CLARITY incorporates a domain adaptation module with cross-lingual and multi-lingual, enabling it to generalize effectively across various platforms and ensuring robust performance even in dynamic online environments. We evaluate CLARITY across multiple benchmark datasets and GPUs, including Kaggle’s Tesla P100, Colab Pro’s NVIDIA T4, and NVIDIA A100. The results demonstrate a significant reduction in inference time, with the A100 achieving an average inference time of 0.85 seconds per instance–over 30% faster than traditional models–while maintaining competitive accuracy.

Read the paper · More papers on PaperTik