The Construction of a Segmented and Part-of-speech Tagged Archaic Chinese Corpus: A Case Study on Huainanzi

Lau Kam · Zhongwen xinxi xuebao · 2013

In this paper,we present a segmented and part-of-speech(POS)tagged Archaic Chinese corpus along with its construction process,which is performed by automatic segmentation and tagging with manual correction as post-processing.We use both Modern and Archaic Chinese labeled data for training word segmenter and POS tagger,which are further improved by domain adaptation techniques,as well as by adding linguistic and morphological features derived from the characteristics of Archaic Chinese language.The experimental results showed the effectiveness of our approach.In particular,the domain adaptation techniques and the added features significantly improve POS tagging performance.During our manual correction,we categorize the errors resulted from the automatic segmentation and POS tagging process,and investigate the sources of those errors.Finally,we give the statistics of the resulted corpus on the distributions of words and POS tags.Our work is a preliminary study that could be easily extended to annotating other Archaic Chinese text,and the resulted corpus is a valuable resource for research on Archaic Chinese language.

Read the paper · More papers on PaperTik