Alignment of Single Source Words and Target Multi-word Units from Parallel Corpus

Limin Du · Zhongwen xinxi xuebao · 2003

Multi-word unit includes steady collocation,multi-word phrase and multi-word term,this paper we provide an algorithm for automatic alignment of single source words and target multi-word units from sentence aligned parallel spoken language corpus.Mutual information has been used to extract multi-word units by many other researchers,but the retrieval results mainly depend on the identification of suitable bigrams for the initiation of the iterative process.This algorithm utilizes normalize mutual information difference and normalize t-scores difference between multi target words correspond to the same single source word to extract the multi-word units,then utilizes the even mutual information and even t-score to align the single source words and target multi-word units.In this algorithm,we have applied the Local Bests algorithm,stopword filter and long-length units preference methods et al.The grading of the lexicon can deduce the number of the incorrect entries in the high level lexicon effectively,which makes the translation lexicon more practicably.

Read the paper · More papers on PaperTik