Accurate Parallel Fragment Extraction from Quasi--Comparable Corpora using Alignment Model and Translation Lexicon

Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi · 2013

Although parallel sentences rarely exist in quasi–comparable corpora, there could be parallel fragments that are also helpful for statistical machine translation (SMT). Previous studies cannot accurately extract parallel fragments from quasi–comparable corpora. To solve this problem, we pro-pose an accurate parallel fragment extrac-tion system that uses an alignment model to locate the parallel fragment candidates, and uses an accurate lexicon filter to iden-tify the truly parallel ones. Experimen-tal results indicate that our system can accurately extract parallel fragments, and our proposed method significantly outper-forms a state–of–the–art approach. Fur-thermore, we investigate the factors that may affect the performance of our system in detail.

Read the paper · More papers on PaperTik