Tibetan Chunking Based on Error-Driven Learning Strategy
Wang Tianhan · Zhongwen xinxi xuebao · 2014
Tibetan chunking is aimed at identifying syntactic constituent in Tibetan sentences to facilitate further analysis of sentences.According to the unique characteristics of Tibetan,the paper puts forward an error-driven learning strategy to identify the chunk boundary based on the description system of Tibetan syntactic functional chunk.The specific idea is as follows:we recognize the chunk boundary using the Conditional Random Fields(CRFs)model at first.Then the recognition result is refined through Transformation-based Error-driven Learning(TBL)method and the CRFs error-driven method.The F values of both methods increase 1.65%and 8.36%,respectively.Finally we combine these two error-driven techniques.In the experiment of the Tibetan corpus which contains 18073 words,the precision,recall and F value achieves 94.1%,94.76%and 94.43%,respectively.