The method of synonyms extraction from unannotated corpus
Alexandr Pak, Sergazy Narynov, Arman Serikuly Zharmagambetov, Шолпан Сагындыкова, Zhanat Kenzhebayeva, Irbulat Turemuratovich · 2015
The structuring of large volumes of e-documents assumes the organization of text on several levels, namely paragraphs, sentences, phrases, words. Methods of lexical paradigms extraction using statistical analysis were developed long ago. In this paper we attempt to move from lexical correlatives to the list of synonyms on various levels of generalization on the basis of local and global contexts' statistics.