Research on extraction of translation equivalents from Chinese-English comparable corpus

Shan Yu-qiu · Computer Engineering and Applications Journal · 2007

This paper reviews the classification of corpora and the history of the research on extraction of translation equivalents from comparable corpus.Based on the basic hypothesis of the extraction of translation equivalents from comparable corpus(namely,there exists a correlation between the context distribution of words which are the translation of each other),this paper adopts the following methods to improve the accuracy of candidates of translation equivalents extracted from comparable corpus:To compute the intersection after the bidirectional extraction of translation equivalents;to calculate the word weight factor TF(iw)*IDF(i),and to utilize the POS information of words in the context.This paper describes the various steps in the experiment of the extraction of translation equivalents from comparable corpus,and conducts analysis on the results from the experiment.The results show that the above methods can improve the accuracy of candidates of translation equivalents extracted from comparable corpus.To round up,the paper puts forward issues required for further research.

Read the paper · More papers on PaperTik