Finding Better Segmentation Granularity for Tibetan-Chinese Bidirectional Neural Machine Translation
Xiaoqiang Han, Weizhao Zhang · 2024
Machine translation refers to the conversion of a source language into a target language by a computer and is an important research direction in natural language processing. Given that different segmentation granularities contain distinct syntactic and semantic features and information, existing research shows that different text segmentation granularities can affect the effectiveness of machine translation. This issue is more pronounced in Tibetan-related translations due to its unique linguistic characteristics. Therefore, in this paper, based on a newly built Tibetan-Chinese parallel corpus of 100,000 pairs, we propose a training method for Tibetan-Chinese bidirectional machine translation using various segmentation granularities, such as characters, syllables and words. The experimental results show that in the Tibetan-Chinese bidirectional machine translation task, using syllable segmentation granularities for Tibetan and Chinese texts yields the best translation performance, achieving a Bilingual Evaluation Understudy(BLEU) score of up to 29.08%, compared to other segmentation strategies.