Multi-Corpus Alignment and Cross-Language Semantic Analysis for Chinese-English Machine Translation

Wenhao Wang · 2025

The continuous evolution of pattern recognition, especially deep learning algorithms, has continuously improved the efficiency of machine translation. This study aims to solve problems of multi-corpus alignment and cross-language semantic analysis in the Chinese-English machine translation and proposes a machine translation framework based on a new dictionary alignment and cross-language semantic analysis. The designed algorithm first realizes paragraph-level corpus alignment through dictionary alignment method. The algorithm uses bilingual dictionaries to calculate the translation probability of sentence pairs and optimizes the alignment effect through evaluation functions. Secondly, this study designs a crosslanguage semantic analysis algorithm. Through the syntactic trees and semantic role annotation, the accurate semantic analysis of sentences is achieved. At the same time, combining Markov tree tense annotation and Transformer model, this study proposes a novel Chinese-English machine translation algorithm. This model effectively solves the tense mismatch problem and improves the translation quality through multi-head attention mechanism. The experiment is verified using Tatoeba Chinese-English data-set. The proposed algorithm achieves the highest BLEU score at different sentence sizes (1000 to 10000). The model also has high robustness. When the sentence size is 1000, the BLEU score of the proposed algorithm is 98.37%, and it remains at$\mathbf{9 4. 6 0 \%}$when the sentence size is$\mathbf{1 0 0 0 0}$.

Read the paper · More papers on PaperTik