Cyberbullying Detection: An Investigation into Natural Language Processing and Machine Learning Techniques
Davy Viriya Chow, Felicia Natania, Oliverio Theophilus Nathanael, Karli Eka Setiawan, Muhammad Fikri Hasani · 2023
Cyberbullying has been a concerning issue ever since the Internet and smartphones became very popular. There are several Cyberbullying types such as videos, images, text, and audio, but the most common occurrence of cyberbullying is through hate speech and offensive language via text message. Making it a suitable task for application of Natural Language Processing (NLP). Social media which includes various platforms like Instagram, Facebook, and Twitter plays a big part in this matter, as it is the medium for communication that has often been the home of hate speech. Many researchers have taken an initiative to generate a tool in detecting cyberbullying. This paper provides a comparative study using several deep learning algorithms such as BERT (Bidirectional Encoder Representations from Transformers), Bi-LSTM (Bidirectional Long Short-Term Memory), and Bi-GRU (Bidirectional Gated Recurrent Unit) in detecting tweets containing cyberbullying. Preprocessing is included in this study, followed by tokenization and embedding. The evaluation of these models focuses on accuracy, precision, recall and F1-Score, which measure their ability to correctly classify different types of cyberbullying from a dataset of tweets. Furthermore, this study also used a confusion matrix as part of the evaluation process. In our study, we found that the BERT model outperformed the BiGRU BiLSTM model in terms of accuracy, achieving about 96%. Followed by Bi-LSTM and Bi-GRU obtained accuracy of 95% and 94% respectively. The result of this comparative analysis indicates all three models exhibit strong performance, while the BERT model had a higher ability to correctly classify cyberbullying instances.