A Method of Construction of the Chinese and English Bilingual Translation Corpus Based on Web Data Mining
Dongfei Liu, Xing Zhou · 2009
The paper introduces a method of construction of the Chinese and English bilingual translation corpus based on web data mining. To collect huge amount of page data by web spider, and identify bilingual web page by a series of complicated purification and analysis process, then analysis the DOM structure of the two page text, we can get the Chinese and English parallel translation corpus and save them to database. As the corpus accumulated by machine automatically, it has higher efficiency, and the translate content come from the internet, original resource is rich and accurate relatively, it can provide a good reference data for translation software.