Text Toxicity Level Detection Using Deep Contextualized Embedding Models

Omar Elgendy, Ali Bou Nassif, Bassel Soudan · Chronicle of computing · 2024

Toxic text is a critical aspect of social media, particularly in today's digital landscape. With the spread of online communication, it has become increasingly easy for individuals to spread harmful or offensive content. Toxic texts include the spread of misinformation, the promotion of hate speech, bullying, and the erosion of trust in online communities. Text toxicity detection algorithms can help to identify and mitigate these negative effects by automatically flagging potentially harmful content. This allows social media platforms to intervene and take appropriate action, such as removing the content or warning the user. Usually, social media platforms offer a reporting strategy which acts after a human decision is made. However, social media now requires an automated system to do this task. In this work, we proposed a Deep learning Regression model to predict the toxicity level in text. Additionally, we fine-tuned multiple Bert models for this task. Our work was evaluated using Mean Square Error, Root Mean Square Error and Mean Absolute Error compared to the testing set of the data and we got for the base model MSE of 0.562, RMSE of 0.750 and MAE of 0.364 but for BERT we got MSE 0.403, RMSE 0.635 and MAE 0.232.

Read the paper · More papers on PaperTik