Checking the Feasibility of the Sejong Bilingual Corpus for Statistical Machine Translation
Sanghoun Song, Francis T. Bond · 언어와 언어학 · 2009
Statistical machine translation research has not made much progress with Korean because of the lack of open Korean resources. This paper proposes using the Sejong Bilingual Corpus as a resource for working on with Korean. The corpus is freely available for research use. Because it covers various genres, a machine translation system based on it is relatively robust across different domains. The corpus can further be combined with other bilingual resources. This paper, the first attempt to apply the Sejong Bilingual Corpus to machine translation on a comprehensive scale, demonstrates that the Sejong Bilingual Corpus can be used as seed data for training a statistical machine translation system.