Context-Aware Models for Text Classification in Sensitive Content Detection

Shashishekhar Ramagundam · International Journal of Scientific Research in Science Engineering and Technology · 2019

Sensitive content detection plays a pivotal role in ensuring the safety and integrity of digital platforms, especially with the increasing volume of user-generated content. Traditional models for content moderation often rely on keyword-based filtering systems that detect explicit offensive terms but fail to identify more subtle forms of harmful content where context plays a significant role. This paper presents a context-aware model for detecting sensitive content that integrates contextual embeddings from transformer-based models like BERT, coupled with deep learning techniques. Our proposed model leverages the power of contextual information, allowing it to understand the nuanced meaning behind text based on its surrounding words and context. The model was evaluated using the Hate Speech Dataset, and our results show a significant improvement in the detection of sensitive content compared to traditional rule-based and keyword-based models. Specifically, the context-aware model achieved a maximum accuracy of 88%, while the baseline rule-based model reached only 70% accuracy. By focusing on context, our approach improves the accuracy, recall, and precision in identifying not only direct hate speech but also more subtle forms of cyberbullying, harassment, and inappropriate language. The proposed method demonstrates the potential of context-aware models in enhancing content moderation, ensuring safer online interactions and contributing to more robust, scalable solutions for sensitive content detection.

Read the paper · More papers on PaperTik