Aligning English-Chinese Words Without Bilingual Dictionary
Lu Xue · Chinese Journal of Computers · 2004
One of the bilingual corpus processing methods is the alignment of two languages on each linguistic level. Much research on word alignment between Indo-European languages has been done before, however, much less has been done on English-Chinese alignment. This paper proposes a corpus-based model for word alignment between English and Chinese. It formalizes natural languages into sets, and the intersection and difference of the sets to implement the word alignment. At the same time, the effect of word order and repetition is considered. The model includes a set of sub-models: minimum intersection model, minimum difference model, hybrid model, mono-directional model, bi-directional model, union model, and surrounding model. The English→Chinese mono-directional model is used to generate 1-m parallels, and the English←Chinese model is used to generate n-1 parallels. The union model and surrounding model are used to generate n-m parallels from the 1-m and n-1 parallels. The intersection of any two generated parallels in a sentence pair is empty, and the parallels themselves are minimum. This method can be used for alignment of both high-frequency words and low-frequency words, and is tolerant with Chinese word segmentation errors and unknown words. The typical characteristic of this model is that it needs few linguistic knowledge and resource. Experimental results show that the larger is the homogeneous corpus scale, the higher precision and recall rate can be obtained.