Detection of cyber harassment (cyberbullying) on Instagram using naïve bayes classifier with bag of words and lexicon based features
Putra Pandu Adikara, Sigit Adinugroho, S Alamia Haque Insani · 2020
Instagram is a very popular social media across the world, with varied users from teenagers to adults. By using Instagram people are able to share photos or videos through social networks. Instagram also provides a lot of features, one of the features is the comment section. However, there are so many Instagram users who make social media as a platform to harass others. Bullying or harassing can affect the psychological condition and in the extreme condition can drive people to suicide. The focus of this research is to detect cyberbullying on Instagram comment into two classes, one is classified as cyberbullying and the other is non-cyberbullying. If we can successfully detect cyberbullying comment, it should help to prevent the cyberbullying act before it happens. The detection process consists of several steps, starts with preprocessing, followed by feature extraction, and the last is classification or in this case, cyberbullying detection. In this research, Naïve Bayes classifier with Bag of Words and Lexicon based features is employed to detect the cyberbullying. The Bag of Words features are extracted from the terms occurred in the comment and Lexicon-based features are extracted by using a dictionary or commonly known as sentiment lexicon. Since Indonesian is a low resource language, it is interesting and challenging to investigate this topic by using Indonesian dataset. In this experiment, the highest evaluation results are obtained by combining Bag of Word features and Lexicon-based features than using the features independently. We use 5-fold cross-validation and the system yields accuracy 0.872, precision 0.948, recall 0.824, and f-measure 0.874.