Corpus of Multi-level Processing for Modern Chinese
Shiwen Yu, Duan, Huiming (Peking University), Yunfang Wu · 2018
Peking University Institute of Computational Linguistics began to research the multi-level processing of the modern Chinese from 1992, and annotated corpus of the People's Daily, 1998 from April 1999 to April 2002. The modern Chinese multi-level processing corpus includes 52 million words of basic processing corpus (word segmentation, part of speech tagging, named entity annotation, phonetic transcription), 28 million words of the same-shaped annotation corpus, in addition, 560,000 words corpus marked parallel structure.