Edit Distance: A New Data Selection Criterion for Domain Adaptation in SMT
Longyue Wang, Derek F. Wong, Lidia Sam Chao, Junwen Xing, Yi Lü, Isabel M. Trancoso · 2013
This paper aims at effective use of training da-ta by extracting sentences from large general-domain corpora to adapt statistical machine translation systems to domain-specific data. We regard this task as a problem of filtering training sentences with respect to the target domain1 via different similarity metrics. Thus, we give new insights into when data selection model can best benefit the in-domain transla-tion. Based on the investigation of the state-of-the-art similarity metrics, we propose edit dis-tance as a new data selection criterion for this topic. To evaluate this proposal, we compare it with other methods on a large dataset. Com-parative experiments are conducted on Chi-nese-English travel dialog domain and the re-sults indicate that the proposed approach achieves a significant improvement over the baseline system (+4.36 BLEU) as well as the best rival model (+1.23 BLEU) using a much smaller training subset. This study may have a significant impact on mining very large corpo-ra in a computationally-limited environment. 1