Translating Between Closely Related Languages in Statistical Machine Translation
Bryce Miller · 2008
Minor languages are gaining more and more status these days, but there is still little parallel data for minority languages, which can be used by a normal SMT system. If we are to translate these languages while still eschewing rule-based systems, then something different must be done with the resources we currently have. The approach was to design and implement a translation model which took advantage of the cross-linguistic correspondences in closely related languages, and the sub-word level. This model was tested against a baseline for accuracy. The model was adjusted by varying word segment sizes, and by varying weightings. Ten language pairs were used. Language pairs translating from Swedish had a marginal improvement above the baseline (0.9 % into Danish, 0.7 % into Norwegian). All other language pairs saw no improvement above the baseline, in any experiment. While the method has not worked for the majority of language pairs, the marginal improvement shown by pairs from Swedish means that this method can work for certain language pairs, and perhaps could work for all, after some improvement. 2