Deep Learning Network Models to Categorize Texts According to Author's Gender and to Identify Text Sentiment
Alexander Sboev, Tatiana Litvinova, Irina Voronina, Гудовских Дмитрий Владимирович, Рыбка Роман Борисович · 2016
In the present article, we consider a problem to evaluate the gain in accuracy of using deep learning network for two language tasks: the automatic text classification according to the authors gender and to identify text sentiment. A preexisting corpus of Russian-language texts RusPersonality labeled with information on their authors (gender, age, psychological testing and so on) has been used for gender task along with the materials of the SentiRuEval competition for evaluating the sentiment of tweets. We have performed the comparative study of machine learning techniques for both tasks on the Russian-language texts. In case of gender tasks the bias in topics and genre was deliberately removed. The obtained neuronet models of deep learning demonstrate accuracy close to the state-of-the-art and even higher: for the gender identification up to 0.86 +/- 0.03 in Accuracy, 0.86 in F1-score, for sentiment classification the best model demonstrates F1 scores with micro average of 0.57 and macro average of 0.61 for banks dataset, and F1-micro of 0.61 and F1-macro of 0.74 for telekom.