The Russian Language Text Corpus for Testing Algorithms of Topic Model

Sergey Nikolaevich Karpovich · Informatics and Automation · 2015

This paper describes the process of creating Russian language text corpus which is specialized for testing algorithms of probabilistic topic model. The articles of Wikinews licensed by Creative Commons Attribution 2.5 Generic (CC BY 2.5) were used as a source of texts for corpus. The stage of text's preprocessing and markup are described in the conclusion. We proposed an original markup of text corpus for testing algorithms of topic modeling.

Read the paper · More papers on PaperTik