Exploring linguistic features for extremist texts detection (on the material of Russian-speaking illegal texts)

Dmitry Alexeevich Devyatkin, Ivan Smirnov, Margarita Ananyeva, Maria Kobozeva, A. M. Chepovskiy, Fyodor Solovyev · 2017

In this paper we present results of a research on automatic extremist text detection. For this purpose an experimental dataset in the Russian language was created. According to the Russian legislation we cannot make it publicly available. We compared various classification methods (multinomial naive Bayes, logistic regression, linear SVM, random forest, and gradient boosting) and evaluated the contribution of differentiating features (lexical, semantic and psycholinguistic) to classification quality. The results of experiments show that psycholinguistic and semantic features are promising for extremist text detection.

Read the paper · More papers on PaperTik