TM‐HOL: Topic memory model for detection of hate speech and offensive language

Jing Chen, Kun Ma, Ke Ji, Zhenxiang Chen · Concurrency and Computation Practice and Experience · 2021

Abstract In the era of the explosion of digital content of large‐scale self‐media, user‐friendly social platforms such as Twitter and Facebook, provide opportunities for people to express their ideas and opinions freely. Due to lack of restrictions, hateful speech and its exposure can have profound psychological impacts on society. Current social networking platform is over‐reliant on the manual check, and it is labor‐intensive and time‐consuming. Although there are many machines learning methods for the detection of hate speech, short text with character limit on social platforms is more challenging for the detection of hate speech and offensive language. To address the problem of data sparsity, we have proposed a topic memory model for hate speech and offensive language detection (abbreviated as TM‐HOL). Potential topics are generated with our encoder and decoder to enrich short text features. Two memory matrices correspond to the topic words and the text, and the hate feature matrix is used to learn the syntactic features. It is demonstrated that our proposed method is effective on three datasets, performing better weighted‐F1.

Read the paper · More papers on PaperTik