DISCRIMINATIVE LEARNING APPROACHES FOR THE STATISTICAL PROCESSING OF CHINESE
Yue Zhang · 2009
We study discriminative approaches to statistical Chinese processing, including word segmentation, pos-tagging and parsing with constituent and dependency grammars. As the main approach of this thesis, we use a global linear model, trained by the generalized perceptron algorithm, together with beam-search decoding, to build our statistical systems. The combination of perceptron training and beam-search decoding leads to highly competitive accuracy and efficiency for all the tasks we investigate. For word segmentation, we propose a word-based approach that achieves competitive accuracy to the best character-based systems, without mapping segmentation into a sequence labeling problem. For pos-tagging, we built a joint system that performs word segmentation and tagging simultaneously, showing that it improves the accuracy over the traditional pipeline approach by enabling information interaction and avoiding error propagation. For constituent parsing, we develop our model based on a shift-reduce parsing algorithm, which currently provides state-of-the-art performance for Chinese, using a global model and beam-search, showing that it achieves comparable accuracy to the best scores in the literature. For dependency parsing, we combine the two predominant statistical methods into a single system using a global linear model, and achieve the current best accuracy for Chinese. One main advantage of the discriminative approach to statistical nlp is the freedom to define features that represent contextual information. For all the problems studied in this thesis, we achieve state-of-the-art accuracies by utilizing a wide range of information in a discriminative model. We conclude that the discriminative method is a competitive choice for the statistical processing of the Chinese language.