Author’s Solutions to PL-EN Corpora Processing Problems

Krzysztof Wołk · 2019

This chapter describes a new method for mining parallel data from comparable corpora using a pipeline of tools. It presents the adaptation of the Bilingual Evaluation Understudy (BLEU) evaluation metric for the needs of Polish-English translation. The chapter discusses enhancements made to the BLEU metric and the evaluation of the enhanced statistical machine translation metric. Non-parallel multilingual data exist in far greater quantities than parallel corpora, but parallel sentences are a much more useful resource. A parser was used for building subject-aligned comparable corpora from Wikipedia articles. The chapter examines the design of the machine translation experiments performed during this research. These experiments involved the machine translation of Translanguage English Database (TED) lectures, subtitles, EuroParl proceedings and medical texts, as well as pruning experiments. Experiments were performed in order to compare the performance of the proposed method with several other sentence alignment implementations on the TED lectures.

Read the paper · More papers on PaperTik