A Method of Merging Corpora in Different Annotation Standards:An Application on Statistics Chinese Lexical Analysis
Jin Xu · Zhongwen xinxi xuebao · 2012
Large scale manually annotated corpora are usually used in research on statistical Chinese lexical analysis.The scale and quality of corpora affect the performance of statistical lexical analysis directly.Corpora in high quality and high rate of coverage are very valuable but limited,and it is very difficult to combine corpora of different domains directly since they are different in segmentation and part of speech(POS) tagging standards.These problems make it difficult to utilize existing resources and prevent the performance improvment in Chinese lexical analysis.To address this issue,this paper presents a simple but effective strategy to optimize the performance and domain adaptability of Chinese lexical analysis by merging different domains' corpora automatically.Our experiments verify the validity,the stronger practicability,and the scalability to multiple corpora of the proposed method.