Domain Adaptation of Statistical Machine Translation using Web-Crawled Resources: A Case Study
Pavel Pecina, Antonio Toral, Vassilis Papavassiliou, Prokopis Prokopidis, Josef van Genabith · 2012
We tackle the problem of domain adapta-tion of Statistical Machine Translation by exploiting domain-specific data acquired by domain-focused web-crawling. We de-sign and evaluate a procedure for auto-matic acquisition of monolingual and par-allel data and their exploitation for train-ing, tuning, and testing in a phrase-based Statistical Machine Translation system. We present a strategy for using such resources depending on their availability and quan-tity supported by results of a large-scale evaluation on the domains of Natural En-vironment and Labour Legislation and two language pairs: English–French, English--Greek. The average observed increase of BLEU is substantial at 49.5 % relative. 1