ALIGNING SENTENCES IN PARALLEL CORPORA USING SELF-EXTRACTED LEXICAL INFORMATION
Sheng Zhu · Chinese Journal of Computers · 1998
Parallel corpora alignment is a key issue in the research of new generation of MT. Thereare two main methods in sentence alignment, i. e., length-based and lexicon-based methods. Thesetwo methods have different characteristics. The former is efficient and easy to implement, but theprecision is not satisfactory, versus the latter. This paper proposes a novel method to alignsentences in Chinese-English parallel corpora. First, the rough result is obtained using thelengthbased method. Then anchors are identified in the texts to reduce the complexity. Some lexicalcorrespondence is also extracted. Finally, the extracted lexical correspondence information is applied infine alignment using lexicon--method. The experimental result shows that this new method cangreatly reduce errors of alignment.