Research on the building of Chinese-Uygur bilingual corpus based on alignment technology

Xuegang Hu · Journal of Hefei University of Technology · 2011

Focusing on the Uygur,this paper proposes an approach to building a Chinese-Uygur bilingual corpus based on alignment technology.According to the characters of both Chinese and Uygur,the Uygur words are segmented,the part of speech(POS) tagging is conducted in light of the table of term frequency,and the alignment of Chinese and Uygur is made to build the Chinese-Uygur bilingual corpus.The above approach is valuable for the research on the Uygur or the other minority languages.

Read the paper · More papers on PaperTik