Cyberbullying Detection and Severity Classification Using Bi-LSTM

K. Nirmala Devi, Vani Rajasekar, P. Jayanthi, Kavin Balasubramani, Kaviya Kandasamy, Keerthana Gowrisankar · 2024

Cyberbullying has grown to be a significant issue in the internet age, affecting individuals of all ages, particularly teens. It involves using internet channels to harass, threaten, or disparage people, often with lingering emotional and psychological repercussions. In social media platforms cyber bullying is becoming more and more common, which has raised the need for efficient tools to identify and remove harmful content. Based on the frequency of objectionable words, a system was created for this research to categorize tweets into four groups: neither, low, medium, or high. A Kaggle dataset was first cleaned and prepared by pre-processing with Keras and Natural Language Toolkit (NLTK). The offending severity level of the tweets was determined using a fuzzy logic with a triangular membership function. Several models were trained using that categorized datasets like Bidirectional Long Short-Term Memory (Bi-LSTM), Bidirectional Encoder Representations from Transformers (BERT-Base) and DistilBERT. Among those models Bi-LSTM showed a high degree of accuracy with a precision score of 95.52%, recall of 95.48%, accuracy score of95.48%, and F1 score of 95.49%.

Read the paper · More papers on PaperTik