Knowledge Translator: Cross-Lingual Course Video Text Style Transform via Imposed Sequential Attention Networks

Jingyi Zhang, Bocheng Zhao, Wenxing Zhang, Qiguang Miao · Electronics · 2025

Massive Online Open Courses (MOOCs) have been growing rapidly in the past few years. Video content is an important carrier for cultural exchange and education popularization, and needs to be translated into multiple language versions to meet the needs of learners from different countries and regions. However, current MOOC video processing solutions rely excessively on manual operations, resulting in low efficiency and difficulty in meeting the urgent requirement for large-scale content translation. Key technical challenges include the accurate localization of embedded text in complex video frames, maintaining style consistency across languages, and preserving text readability and visual quality during translation. Existing methods often struggle with handling diverse text styles, background interference, and language-specific typographic variations. In view of this, this paper proposes an innovative cross-language style transfer algorithm that integrates advanced techniques such as attention mechanisms, latent space mapping, and adaptive instance normalization. Specifically, the algorithm first utilizes attention mechanisms to accurately locate the position of each text in the image, ensuring that subsequent processing can be targeted at specific text areas. Subsequently, by extracting features corresponding to this location information, the algorithm can ensure accurate matching of styles and text features, achieving an effective style transfer. Additionally, this paper introduces a new color loss function aimed at ensuring the consistency of text colors before and after style transfer, further enhancing the visual quality of edited images. Through extensive experimental verification, the algorithm proposed in this paper demonstrated excellent performance on both synthetic and real-world datasets. Compared with existing methods, the algorithm exhibited significant advantages in multiple image evaluation metrics, and the proposed method achieved a 2% improvement in the FID metric and a 20% improvement in the IS metric on relevant datasets compared to SOTA methods. Additionally, both the proposed method and the introduced dataset, PTTEXT, will be made publicly available upon the acceptance of the paper. For additional details, please refer to the project URL, which will be made public after the paper has been accepted.

Read the paper · More papers on PaperTik