Strategies for Optimizing English Translation Models through Reinforcement Learning

Chunxia Zhang, Zheng Li, Fujie Zhai, Jie Gao, Tianyu Wang · 2024

To address the shortcomings of traditional English translation models in semantic accuracy and contextual understanding, this study adopts Proximal Policy Optimization (PPO) to improve translation quality and efficiency by defining accurate reward signals and optimizing model parameter adjustments. The paper first collects a large number of English translation samples, including positive and negative samples, to construct sorting sample data. Then, a feedback model is trained based on the sorting sample data so that it can score positive samples higher than negative samples. Then, based on the scores of the generated translation samples given by the feedback model, PPO is used to fine-tune the translation model. After each iteration, it is determined whether the performance of the translation model meets the preset conditions, otherwise the iteration continues. Finally, the term translation accuracy information is introduced into the translation model to enhance the term translation accuracy and translation with strong domain style, so that the generated sentences are more in line with the characteristics and style of the vertical field. Experimental results show that the optimized model has made significant progress in translation accuracy. The average METEOR (Metric for Evaluation of Translation with Explicit ORdering) score has increased from 0.74 to 0.88, and the average translation time has been reduced from 368.32 milliseconds to 254.88 milliseconds, a reduction of 113.44 milliseconds. This study provides strong technical support for meeting higher translation needs, and has important practical significance for promoting cross-cultural communication and improving the efficiency of international cooperation.

Read the paper · More papers on PaperTik