Identifying Abusive Comments in Hebrew Facebook
Chaya Liebeskind, Shmuel Liebeskind · 2018
In this study, we aim to classify comments as abusive or non-abusive. We develop a Hebrew corpus of user comments annotated for abusive language. Then, we investigate highly sparse n-grams representations as well as denser character n-grams representations for comment abuse classification. Since the comments in social media are usually short, we also investigate four dimension reduction methods, which produce word vectors that collapse similar words into groups. We show that the character n-grams representations outperform all the other representation for the task of identifying abusive comments.