Analyses of Hate and Non-Hate Expressions during Election using NLP

Nico T. Solitana, Charibeth Cheng · 2021

Social media platforms such as Twitter were able to connect people around the world closer by enabling them to actively connect and communicate as well as exchange ideas on social issues; however, it has also become niche for misinformation and hate speech which does not only affects peoples' perception of truth but can also lead to hate crimes as studies confirmed. Recent works in text classification has proved pivotal in detecting hate speech but is continued to be limited by some problems. Given this, the study aimed to further analyze the composition of hate expressions using the collection of tweets posted during the 2016 Philippine Presidential Election. Word bubbles were generated to identify word similarity and differences between hate and non-hate. Principal Component Analysis with K-means clustering was used to further confirm clustering of hate targets against the grouping from annotation. Results showed that both hate, and non-hate tweets have around 55% different terms while the clustering of hate targets based on annotation does not complement the groupings generated in terms of nearest neighbors. Lastly, lexical diversity reveals significant difference between hate classes and targets which suggest that it can be potentially used as a feature to improve hate speech classification.

Read the paper · More papers on PaperTik