English-chinese Parallel Corpora Based on the Automatic Extraction of Terms Dictionary

Liang Ming · Computer Knowledge and Technology · 2009

In the field of natural language processing,the importance of bilingual parallel corpus is increasing.In recent years,many research institutions at home and abroad are building bilingual corpus,and many of the bilingual corpus researchers conducted extensive research.Sentence alignment is an important component of bilingual corpus building,and also the basis work of the machine translation.This paper describes the research background and current situation of the terminology extraction based on the bilingual parallel corpus,and then intro-duced several ways and basic principles used in the sentence alignment.The bilingual corpus the experiment used is already good alignment,which after a Chinese word processing.Use English and Chinese POS tagging tools to tag the Chinese and English Corpus respectively.The term candidate set is produced by statistical the nouns and noun phrases of both corpus.Then translation probability between every English candidate term and its Chinese translation term are calculated.By setting the threshold to filter out some candidates with the English word unrelated to the Chinese translation,finally,select the greatest probability of the English word as a candidate of the Chinese translation of the word by greedy algorithm.

Read the paper · More papers on PaperTik