A corpus of Chinese-Tibetan phrase translation based on semantic enrichment—SECT
Run CHANG, Bo CHEN, Xiaobing ZHAO · China Scientific Data · 2024
Machine translation plays an important role in natural language processing, and increasingly important for promoting political, economic, and cultural exchanges. In high-resource languages like Chinese and English, machine translation has almost reached the accuracy of human translation. However, for low-resource languages like Tibetan, the accuracy of Tibetan machine translation still needs improvement due to the lack of large-scale publicly available parallel corpora. In Chinese-Tibetan translation, when it comes to phrase translation, the existing machine translation results are often inaccurate due to their brevity and the deep semantic information between the lines, such as abbreviations. To improve the capability of translation models to better capture and convey semantic information, this paper constructs a Chinese-Tibetan phrase translation corpus based on semantic information enrichment. This corpus contains 7,000 entries of Chinese-Tibetan phrase pairs. The original data for Chinese-Tibetan phrases is sourced from the Tibetan Language and Writing Network of Tibet, and the enriched semantic information includes Chinese definitions for Chinese phrases and example sentences incorporating the target phrases. This part of the content is obtained through the generation of large language models and professional proofreading. The publication of this dataset is of great value in promoting the development of Chinese-Tibetan information processing.