Ancient Sentence Search Based on Sentence AutoAlignment in Parallel Corpus of Ancient and Modern Chinese
Min Liao · Zhongwen xinxi xuebao · 2008
Along with the Corpus Linguistics' prosperity and development,the research on Example Based Machine Translation(EBMT) has a flourishing prospect.In this area,two problems must be solved: 1) Constructing a large-scale parallel corpus with high accuracy and speed.2) Searching the most similar sentence with the input sentence from the huge aligned examples.This paper aimed at EBMT between ancient and modern Chinese.First,a new translation model was built which takes the length of the sentence,character information and punctuation into account at the same time.Then,a new approach for aligning bilingual sentences automatically was proposed based on genetic algorithm and Dynamic Programming.Finally,a new similarity method was given based on Chinese characters' information entropy.Experimental results showed that our methods achieved good performance.