Extraction of Translation Equivalent Pairs fromChinese-English Parallel Corpus

Chang Baoba · Terminology Standardization & Information Technology · 2002

More and more researchers have recognized the potential value of the parallel corpus in the research on Machine Translation and Machine Aided Transl ation. This paper examines how the translation equivalent pairs could be extract ed from parallel corpus. An iterative algorithm based on degree of word associat ion is proposed to identify the multiword units for Chinese and English. Then a hypothesis-testing approach is used to extract the Chinese-English Translation Equivalent Pairs. We also made comparison between different statistical associa tion measurement and proposed to use categorical hypothesis to improve the perfo rmance of extraction.

Read the paper · More papers on PaperTik