Tibetan Automatic Word Segmentation Method Based on Deep Learning
Gadeng Luosang, Nuo Qun · 2022
The automatic word segmentation of Tibetan is a fundamental task in Tibetan natural language processing. The result of automatic word segmentation will directly affect the other task effect, such as Information Retrieval, Intelligent Question Answering systems (IQAS), and Machine Translation. Current automatic word segmentation uses Tibetan syllables as input, but it is impossible to mine syllable component information. This paper proposes a deep learning model that fuses syllable components, called TBiLSTM. After comparing with the existing Tibetan automatic word segmentation methods, this method has a better effect.