Twitter-based Polarised Embeddings for Abusive Language Detection
Leon Graumas, Roy David, Tommaso Caselli · 2019
We present a method to generate polarised word embeddings using controversial topics as search terms in Twitter as proxies for interactions among social media communities that may be liable to use abusive language. We investigate to what extent models trained with these embeddings perform with respect to generic embeddings across four data sets of abusive language, both in the same domain and out of domain, using simple linear classifiers. Our results show that the polarised embeddings are competitive in the same domain data sets, and perform better in out of domain one.