Tibetan Text Classification Algorithm Based on Syllables

Xiang‐He Meng, Hongzhi Yu, Hui Cao · 2020

Tibetan text classification is one of the core technologies in the field of Tibetan information processing. With the rapid development of the Internet, a large amount of Tibetan Internet text data will be generated every day. Text classification technology can quickly and accurately obtain the required information to solve the problem of out-of-order in text. Tibetan syllables are the basic components of Tibetan text, and each syllable in Tibetan is divided by syllable nodes. This paper proposes a Tibetan syllable as a text representation feature, and uses deep neural network models such as CNN, BiL STM and RCNN to classify Tibetan text. Experiments show that this method has achieved prefect results in different depth neural network classification models.

Read the paper · More papers on PaperTik