Automatic assignment of labels in Topic Modelling for Russian Corpora
Aliya Mirzagitova, Olga Aleksandrovna Mitrofanova · ExLing Conferences · 2016
The main goal of this paper was to improve topic modelling algorithms by introducing automatic topic labelling, a procedure which chooses a label for a cluster of words in a topic. Topic modelling is a widely used statistical technique which allows to reveal internal conceptual organization of text corpora. We have chosen an unsupervised graph-based method and elaborated it with regard to Russian. The proposed algorithm consists of two stages: candidate generation by means of PageRank and morphological filters, and candidate ranking. Our topic labelling experiments on a corpus of encyclopedic texts on linguistics has shown the advantages of labelled topic models for NLP applications.