An Analysis of COVID-19 related Twitter Data for Asian Hate Speech Using Machine Learning Algorithms
Sandeep Shah, Xiaohong Yuan, Zanetta Tyler · 2022
Many people have suffered in this global pandemic of Covid-19 since March 2020. Some suffered from covid-19 disease whereas some suffered from hatred. There are posts and comments on Twitter that reflected Asian hatred blaming them for the coronavirus. Many people expressed their anger and racism through social media and even physically. It is an important task to understand the sentiments of public data to reduce hatred and racism through social media. Many machine learning algorithms can be used to classify Twitter data. In this research, we use two different machine learning classifiers to classify COVID-related Twitter data: Support Vector Machine (SVM) and Random Forest. We compare the performance of both learning methods according to their prediction accuracy, precision, recall, and F1-score. We then predict the tweet label using the data for each month from April to November 2020 using the SVM model. The label used in the dataset are Hate, Counter-hate and Neutral. We analyze the ratio of hate and counter-hate tweets and discuss their possible correlation with events that occurred in those months.