The Construction of English-Chinese Parallel Corpus of Medical Works Based on Self-Coded Python Programs

Xiaoxiao Chen, Shili Ge · Procedia Engineering · 2011

In order to provide sufficient training data for statistical machine(-aided) translation in medical field, a large scale English-Chinese parallel corpus of medical works is constructed. Eighteen English medical printed books with Chinese translation are selected as raw materials. With the help of an OCR scanner, all texts are recognized, manually proofread and stored in electrical form. Within a rigid scheme of corpus construction and with the help of a self- coded Python program, English and Chinese texts are separated, sentence aligned and XML marked. After careful manual proofreading, an Internet-based corpus retrieval platform is constructed. The present parallel corpus contains 54,522 sentence pairs and more than 2,500,000 English words / Chinese characters, which can be preliminarily applied in the training and testing of statistical machine(-aided) translation researches in medical field.

Read the paper · More papers on PaperTik