Which is More Effective for Chinese Lexical Analysis via Character Tagging:Above-context Versus Below-context
Fan Xiao-zhong · 2012
Chinese lexical analysis is a foundational task for Chinese information processing.At the current,the mainstream technology of Chinese lexical analysis is based on statistical methods.These methods treat the analysis process as a sequence data tagging problem.Context is the necessary resource not only for obtaining linguistic knowledge in statistical linguistics but also for solving the problem in natural language processing.Chinese lexical analysis needs the help of correlative context.However,are above and below the same important? To overcome the lack of giving the result by the subjective experience,we studied the contribution of above and below for character-based tagging Chinese lexical analysis via the large number of experiments about word segmentation,POS tagging and named entity recognition.Closed evaluations were performed on many kinds of corpus from the international Chinese language processing Bakeoff,and comparative experiments were performed on different feature templates which describe above-context and below-context.Experimental results show that the performance by the below-context increases 6 percentage points than by the above-context.