Research on terms extraction in literature technical based on parallel corpus

Zhong Yu-feng · Journal of Heilongjiang Institute of Technology · 2011

The importance and the distribution circumstances of literature technical terms are introduced first.Extraction methods mostly in use of literature technical terms are sumed up.And then an algorithm for the automatic extraction of bilingual term from English-Chinese parallel corpus is proposed in the paper.Parallel corpus is aligned chiefly by improved statistical method,which is based on character length,and tagged with their part-of-speech categories respectively.The term candidate set is produced by counting the nouns and noun phrases of both corpora.Then the translation probability between every English candidate term and its Chinese translation are calculated.Finally,the experiments of term extraction on Parallel Corpus of Regulations for the Implementation of the Copyright Law of the PRC had been done.

Read the paper · More papers on PaperTik