Research on the Automatic Identification of Tibetan Sentence Boundaries with Maximum Entropy Classifier

Zangtai Cai · Computer Engineering and Science · 2012

The boundary Ientification of Tibetan sentence is the basical research of Tibetan text analysis.It is the essential work to build a Parallel Corpora between Tibetan and other languages,and also it is the base to do Tibetan-Chinese machine translation.The article raises the ways of Boundary Identification of Tibetan sentences through the analyze of the ending forms of Tibetan sentences and the study of it's boundary rules.The method is firstly using the special rules and word forms to identify Tibetan Sentences,and then to make a further identification for those ambiguous sentences by using Maximum Entropy Model.So it can improve the boundary identification rate of Tibetan sentences.

Read the paper · More papers on PaperTik