Research on dpo handwritten text generation algorithm based on fine-grained improvement

Qingting Liu, Shuying Zhao, Junjie Cai · 2025

Sign language text generation is an important research direction within sign language generation systems. The Direct Preference Optimization (DPO) algorithm is a reinforcement learning model based on human feedback. It offers advantages such as controllable outputs, better alignment with human preferences, and independence from reward systems for evaluation. However, DPO-based sign language text generation faces three key challenges: (1) The scarcity of high-quality negative sample datasets in the field of sign language poses significant challenges to model training; therefore, this paper proposes a method leveraging word vector similarity to construct human preference datasets that make full use of negative samples. (2) The global optimization approach of traditional DPO algorithms, which struggles to handle local errors; therefore, a token-level dynamic weighting mechanism is introduced for finer-grained optimization. (3) The lack of systematic evaluation methods for assessing the accuracy and applicability of generated text; therefore, a multi-dimensional evaluation method based on a large model for handwritten text is proposed. Experimental results demonstrate that these fine-grained improvements lead to significant enhancements in the model performance for handwritten text generation tasks, with notable improvements in accuracy, fluency, and contextual understanding. This research provides new directions for advancing sign language text generation technologies.

Read the paper · More papers on PaperTik