Building English Chinese Statistical Translation Models from Semi structured Parallel Texts

Nie Jian · Zhongwen xinxi xuebao · 2001

A statistical translation model tries to capture translation relationships from a set of parallel texts (or translation examples).This paper describes our attempt to train such translation models from a set of semi structured parallel texts in Chinese and English.These texts are gathered from the Web by an automatic mining tool PTMiner.Our work takes advantage of the HTML structure of the texts.Some special processing is necessary on Chinese.Our experiments show that we can obtain a translation precision of about 80% with the trained model.This performance is reasonable for less critical tasks such as cross language information retrieval.This work shows that it is possible to construct a means of query translation at a much lower cost than a machine translation system.

Read the paper · More papers on PaperTik