A System That Analyzes Bengali Text on Facebook Posts Using Machine Learning to Spot Suspicious Content

Meherun Nesa Shraboni, Mohammad Shamsul Arefin · 2025

In the digital landscape, social media platforms have become hotspots for disseminating harmful content, such as political propaganda, cyberbullying, and incitement of violence. To combat this issue, we have developed a machine learning-based system that automatically detects and flags suspicious content in Bengali text uploaded on Facebook. Leveraging supervised learning on a dataset of Bengali text, our system proficiently identifies dubious information. By incorporating AI technology, we can filter out questionable content from many posts on specific websites featuring Bengali text. The system is complemented by a monitoring dashboard, providing robust control over the material and ensuring efficient management and filtration of suspicious content. Our work involves developing a Bengali Text Detector automated system that uses essential domains such as machine learning, natural language processing, social media, hate speech, false information, and social issues. Our work encompasses developing an automated system called Bengali Text Detector that uses essential domains such as machine learning, natural language processing, social media, hate speech, false information, and social issues introducing the pioneering Bangla Suspicious Text Dataset (STD) 1 , which includes words and lines. Derived from Bengali comments posted in a social media group, it serves as ground truth texts. The dataset comprises 19,629 Bengali text data collected from various sources. We analyze and classify the text into “suspicious” and “unsuspicious” categories, achieving an impressive 70% accuracy. Our system effectively identifies and mitigates the spread of misleading information, hate speech, and harmful content on social networking platforms.

Read the paper · More papers on PaperTik