Construction of Machine Translation Model Based on Multimodal Information Fusion Algorithm and its Effect Analysis
Yan Zhang · Procedia Computer Science · 2025
Traditional machine translation (MT) is difficult to effectively handle complex syntactic structures and semantic information. Some issues that lead to poor translation quality include a lack of multilingual data and ineffective feature representation. The paper constructs a translation model based on multimodal information fusion algorithm to improve the quality of MT. In order to more comprehensively capture the contextual features of the source language vocabulary, the compound words that are difficult to segment directly are decomposed into affixes, which are refined into smaller language units and converted into vector form through word embedding technology. The Long Short-Term Memory (LSTM) and Multi-Head Self-Attention (MHSA) mechanism are used to build a MT model, namely the LSTM-MHSA model, which uses a bidirectional LSTM as the basic structure of the encoder and integrates multimodal information to encode the source language material; In the decoding stage, MHSA is introduced to dynamically adjust the influence of different modalities, so that the model can accurately measure the importance of each modality based on multimodal input information, thereby generating more accurate and authentic translation text. Experiments show that the BLEU (Bilingual Evaluation Understudy), TER (Translation Edit Rate) and English translation accuracy of the LSTM-MHSA model are 63.4%, 23.1% and 95.4% respectively. The MT model examined in this study has improved translation efficiency.