Toxic Speech Classification using Machine Learning Algorithms
Pabba Sumanth, Syed Samiuddin, Kovelakuntla Jamal, Srikanth Domakonda, Pathi Shivani · 2022
In today's era of online social media platforms, there has been a massive surge in the propagation of toxic content speech. They provide many betterments. However, persons with considerable differences in their viewpoints have contributed to an increase in lethality of people in internet posts and debates.. With the outbreak of the pandemic, corporations, educational institutions, students, and the general public have all increased their usage in web sites. For a long time, the growing popularity of internet platforms like Twitter and Facebook has been a major cause of anxiety. These platforms not only allow for improved communication, but they also allow the users to express their thoughts, which are quickly shared with the rest of the world. Furthermore, given the diversity of these platforms' users' histories, beliefs, race, and customs, many of them choose to use disparaging, abusive, and antagonistic language while interacting with those who do not share their background. This online toxicity has been increasing exponentially by advancements provided by these social media platforms in this emerging world under the cloud of anonymity. Unlike manually, this problem can be solved using Machine Learning. Phrases like “Obscene”,”Toxic”, “Severe Toxic”,”Threat”, “Insult”,”Identity Hate” are used mutually and hence have been incorporated under “Toxic” speech content. As a result, it is vital to recognise and eliminate toxic speech from internet - based social media networks naturally. The numerous varieties of Machine Learning approaches, such as traditional Machine Learning, ensemble approach are explored in this paper. We use a corpus collected from online platform twitter to do binary and multi-class classification and investigate two techniques.: (a) a method which consists in extracting of word embeddings and then generating the model; (b)Improving the existing models- RF, DT, VC, LR, KNN. Any other sort of social media comment can be analyzed using the proposed methods. By this, we developed a model that can classify given comments into different categories of toxicity with greater precision, recall, and accuracy score.