Binary-class and Multi-class Chinese Textural Entailment System Description in NTCIR-9 RITE.
Shih-Hung Wu, Wan-Chi Huang, Liang-Pu Chen, Tsun Ku · NTCIR · 2011
paper, we describe the details of our system for NTCIR-9 RITE. We sent 3 runs for each of the four sub-tasks: CT-BC, CT- MC, CS-BC, and CS-MC. Our approach to the NTCIR-9 RITE task is based on the standard supervised learning classification. We integrate available computational linguistic resources of Chinese language processing to build the system in a statistical natural language processing approach. First, we observed the training corpus and list all possible features. Second, we test the features on training data and find features that can be used to identify textual entailment. The features include surface text, semantic and syntactical information, such as POS tagging, NER tagging, and dependency relation. An automatic annotation subsystem is built to annotate the training corpus. Finally, the annotated data is used in training statistical models and build the classifier for the RITE 1 subtasks.