Refining semi-automatic parallel corpus creation for Zulu to English statistical machine translation
Gideon Kotzé · 2016
Although their use in training quality machine translation systems has been proven, parallel corpora - large collections of translated texts - are generally hard to come by for the majority of languages. To counteract this fact, a relatively small collection may be processed in more depth by further cleaning and more accurately splitting and aligning the texts. We apply this to an existing English/Zulu parallel corpus that has been used for statistical machine translation experiments. After these preprocessing steps, we run the same experiments for comparative purposes. Our results suggest that compatibility of bitexts, the choice of sentence splitters used on different parts of the text, as well as manual work, may have a notable effect on both the corpus size and on automatic translation quality.