Patent SMT Based on Combined Phrases for NTCIR-7

Junguo Zhu, Haoliang Qi, Muyun Yang, Jufeng Li, Sheng Li · 2008

In this paper, we describe a combined phrase approach to the Statistical Machine Translation of Japanese patents into English. To resolve the segmentation errors caused by the rich OOV (out-of-vocabulary) words in the patent texts, the character based translation phrases are first employed. Then the word based translation phrases are established to utilize the dependable word level information. Finally the two translation phrases tables are linearly combined to capture both character and word level translation correspondences. Preliminary experiments on NTCIR-7 corpus indicate that the BLEU scores of the proposed method significantly out-perform the usual word based approach.

Read the paper · More papers on PaperTik