Exploiting Parallel Corpus for Handling Out-of-vocabulary Words
Juan Luo, John F. Tinsley, John Tinsley, 51576, Yves Lepage, 51279, 70573608 · Institutional Repositories DataBase (IRDB) · 2013
This paper presents a hybrid model for han-dling out-of-vocabulary words in Japanese-to-English statistical machine translation output by exploiting parallel corpus. As the Japanese writing system makes use of four different script sets (kanji, hira-gana, katakana, and romaji), we treat these scripts differently. A machine translitera-tion model is built to transliterate out-of-vocabulary Japanese katakana words into English words. A Japanese dependency structure analyzer is employed to tackle out-of-vocabulary kanji and hiragana words. The evaluation results demonstrate that it is an effective approach for addressing out-of-vocabulary word problems and decreasing the OOVs rate in the Japanese-to-English machine translation tasks. 1