Aggression Detection in Social Media Texts using Machine Learning and Deep Learning Models
Fatema Huseni Vadnagarwala, Avinash J. Agrawal · 2024
In recent times with profuse usage of social media for communication and emotional content sharing, the need for effective control measures is required to prevent irrelevant and aggressive messages. The impact of aggressive and offensive tweets, post and messages on the cyber platform is negative. Our study focuses on how to identify such content and proposes the elaborate study of performance of multiple Machine learning algorithms (such as Naïve Bayes, Random Forest, Gradient Boost, Logistic regression and XGBoost) and Deep Learning algorithms (such as LSTM, BiLSTM and DNN), along with execution of latest BERT models with finetuning approach implemented to give a comparative analysis. The main focus of this study is to analyse the impact of the quality of dataset and the performance of the models. It gives an intuition of achieving better results for multilingual, unstructured and non-contextual lexicon data classification. It analyses the hidden masking of aggressive words with symbols and gives improved result for prevention of bypassing the detection algorithms. In the end we propose an ensemble mechanism for collectively using the best performing algorithms in a stacking classifier for improved classification. Our study focuses on performance improvement without any manual intervention and manual feature extraction. The results achieved is with F 1-score 94% approx.