Research on Sentence Segmentation and Punctuation in Ancient Chinese

Han Cai-hua · Journal of Henan University · 2009

Data sparseness is a primary challenge in sentence segmentation and punctuation in ancient Chinese using natural language processing technology.In order to overcome this difficulty,a 6-tag set was designed and a method based on cascaded Conditional Random Fields was proposed.The main idea is as follows: based on the 6-tag set,a low level model determines the boundaries of sentences according to observation sequence and a high level model punctuates sentences taking consideration of both observation sequence and low level's results.Close test and open test were done based on approximate 5M mixed corpus respectively.The F measure of sentence segmentation and punctuation are 96.48% and 91.35% respectively in close test,and those are 71.42% and 67.67% respectively in open test.

Read the paper · More papers on PaperTik