Research on Optimization of Korean Translation Model Based on Deep Learning
Xinfeng Wang, Debin Han · 2024
This paper studies the optimization of Korean translation model based on deep learning (DL). In view of the unique grammatical structure and rich vocabulary changes in Korean, this paper first constructs a rich training data set by using multiple open parallel corpora and supplementing specific domain data through web crawler and manual translation. Then, the Transformer model is adopted as the infrastructure, and the morphological analyzer and Byte-Pair Encoding (BPE) technology are introduced to deal with the lexical morphology and unknown words in Korean. By designing several groups of comparative experiments, including baseline model, model for optimizing Korean characteristics, model for enhancing the diversity of data sets and model for pre-training and fine-tuning, the influence of different optimization strategies on translation quality is evaluated. The experimental results show that the optimized model has significantly improved the evaluation indexes such as BLEU, METEOR and NIST, and the translation quality, fluency and naturalness have been significantly improved. Specifically, the optimized model 3 performed best in all evaluation indicators, with an average increase of 6% for BLEU Score, 8% for METEOR Score and 17% for NIST Score. This study not only provides an effective strategy for optimizing the Korean translation model, but also provides a useful reference for multilingual translation research.