NLP-Driven Translation Quality Evaluation: A Joint Optimization Approach for Prompt Design and Post-Editing
Yameng Pan, Yi Huang, Jiaxin Lin · 2025
Large language models generate content by learning from vast amounts of text data. By optimizing prompts, users can communicate their needs to the model more clearly. This study focuses on the translation of Jin Gui Yao Lue, comparing the human translation with translations produced by ChatGPT-4o and DeepL. Based on a self-constructed parallel corpus and multidimensional indexes, the findings reveal: (1) Translations of ChatGPT and Luo show similar lexical features, while DeepL, constrained by source language structures, exhibits higher lexical repetition; (2) Luo's translation favors complex noun phrases to convey TCM concepts, ChatGPT prefers verb-driven simplified structures, and DeepL retains the paratactic and lengthy sentence style typical of classical Chinese; (3) ChatGPT relies more on implicit logical cohesion with higher information density, whereas DeepL presents greater interactivity due to its formal equivalence strategy. Incorporating descriptions of typical linguistic flaws found in ChatGPT-generated translations into the prompt led to significant improvements in both BLEU and METEOR scores, confirming a substantial enhancement in translation quality. This study provides clear directions for optimizing ChatGPT-based TCM translation and empirical evidence to support the development of AI-based machine translation post-editing (MTPE) strategies.