Extraction of Bilingual Technical Terms for Chinese-Japanese Patent Translation
Wei Ping Yang, Jinghui Yan, Yves Lepage · 2016
The translation of patents or scientific papers is a key issue that should be helped by the use of statistical machine translation (SMT).In this paper, we propose a method to improve Chinese-Japanese patent SMT by premarking the training corpus with aligned bilingual multi-word terms.We automatically extract multi-word terms from monolingual corpora by combining statistical and linguistic filtering methods.We use the sampling-based alignment method to identify aligned terms and set some threshold on translation probabilities to select the most promising bilingual multi-word terms.We pre-mark a Chinese-Japanese training corpus with such selected aligned bilingual multi-word terms.We obtain the performance of over 70% precision in bilingual term extraction and a significant improvement of BLEU scores in our experiments on a Chinese-Japanese patent parallel corpus. IntroductionChina and Japan are producing a large amount of scientific journals and patents in their respective languages.The World Intellectual Property Organization (WIPO) Indicators 1 show that China was the first country for patent applications in 2013.Japan was the first country for patent grants in 2013.Much of current scientific development in China or Japan is not readily available to non-Chinese or non-Japanese speaking scientists.Additionally, China and Japan are more efficient at converting research