Automated Hate Speech Detection on Twitter

Garima Koushik, K. Raja Rajeswari, Suresh Kannan Muthusamy · 2019

With the sudden increase in micro-blogging websites such as Twitter, Facebook, and Tumbler, the communication between people becomes indirect and reliable; people from different educational backgrounds, cultures share their opinion on different aspects of life every day. This has resulted in conflicts among people. As a result, the use of hate speech becomes a very serious problem. Manual detection of such content from these websites is a very tedious task. Hate speech is the use of aggressive, violent or offensive language which targets a specific group of people sharing common property, this property can be their gender, ethnic group or their believes and regions. The proposed model is capable to detect hate content on Twitter automatically. This approach is based on a bag of words and TFIDF (term frequency-inverse document frequency) approach. These features are used to train machine learning classifiers. Exhaustive experiments are conducted on existing twitter dataset and the accuracy obtained by logistic regression classifier is equal to 94.11% on detecting whether a particular tweet is hateful or not.

Read the paper · More papers on PaperTik