Hate Speech Detection of Roman Urdu Using Transformer Models

Rehman Ullah, Waseem Ullah Khan, Safdar Nawaz Khan Marwat, Jasim Ullah, Fatima Irshad · 2024

Hate speech means that giving discrimination about someone, such as a person, a group or any community on the base of religion, personality of someone, race, sexual harassment, gender, or on the base of other identities. To prevent hated speeches different laws has been made by the governments of different countries. Which has worked a lot for the citizen's safety of these countries and their citizens are safe from such hated speeches up to some extent. Too many researchers and scientists has worked a lot in NLP to prevent hated speeches throughout the world, and they been successes up to some extent in their aim. As we all see nowadays on different platforms of social media that hated speeches are detected. We have used some of transformer methods and some of non-transformer methods, such as BERT Bidirectional Encoder Representations from Transformers (BERT) in transformer method and DISTEL BERT in transformer method and transformer model in transformer method. And we have LSTM, LSTM + attention layer, CNN and LSTM in nontransformer method. And we have trained it by one of the roman Urdu datasets which consists 30000 thousand of roman (Urdu) tweets, moreover we will check or look for accuracy precession, F1 score and recall in it. Then we compared those models in which BERT has the best result that's transformer model.

Read the paper · More papers on PaperTik