Method study of deeper processing for Tibetan corpus

Zangtai Cai · Computer Engineering and Applications Journal · 2012

As the constant development and improvement of natural language information processing,enormous linguistic material text processing has become a hot topic in the area of computational linguistics.One important reason is that it can collect the demanding knowledge from the huge corpus.This article puts together the development experience of the 973 project—— Studies on syncopate-dimensional norms of the Tibetan corpus,elaborates on the large-scale construction of the Banzhiada Tibetan corpus,the design and the realization of the syncopate-dimensional dictionary storehouse and the syncopate-dimensional software.It mainly discusses the index structure and the lookup algorithm of the dictionary storehouse,the matching algorithm case auxiliary words block and the decompression algorithm of syncopate-dimensional software.

Read the paper · More papers on PaperTik