Probabilistic methods of automated dynamic thesauri creation from heterogeneous knowledge sources

Andrey S. Pisarev · 2017

Probabilistic methods of the automated formation of dynamic thesauri are considered on the basis of algorithms for processing, analyzing and classifying linguistic resources of a very large volume from heterogeneous data sources in application to the support systems of learning processes. Probabilistic methods are implemented in the automated software tool “Ontomaster-Ontology”. An essential feature of the implementation is the provision of automated processing of very large volumes of linguistic resources from heterogeneous sources. Probabilistic methods have been applied to the formation of thematic text corpora in Russian and English languages, extraction of terms of subject domains, classification of documents. Reprezented the results of the study of improving the efficiency of document classification based on the Bayesian approach are presented using the n-gram model and verbose terms. Probabilistic methods of computer linguistics are applied in the development of a system for supporting the process of teaching students in the direction “Information Systems and Technologies”.

Read the paper · More papers on PaperTik