SET EXPANSION OF CONTEXTUAL SEMANTIC RELATIONS: AN ALTERNATIVE TO FULL CORPUS ANNOTATION FOR SUPERVISED CLASSIFICATION
André Kenji Horie, Mitsuru Ishizuka · International Journal of Semantic Computing · 2012
Recent approaches for classification of semantic relations are based on supervised learning using large training datasets. Due to the high cost of annotating such data and to the class imbalance problem, alternatives for minimizing the effort of full corpus annotation are required. In set expansion, one of such alternatives, given a small initial training set, new relevant instances are acquired from a large corpus. However, when dealing with contextual semantic relations, which are relations that are highly dependent on the context within the sentence, set expansion is not trivial, since instances are not directly queryable and filtering requires classification under a very restricted number of training instances. This work thus proposes a bootstrapped set expansion method for contextual semantic relations. It performs a best effort extraction using the Web, and a two-stage filtering of candidate instances, the first based on syntactic patterns and the second using a feature distance-based classifier designed for the low frequency setting. The relevance of the output is measured experimentally by using the expanded set as the training data of the supervised classification task, observing an incremental improvement in performance after each bootstrapping iteration when compared to values using the unexpanded training data.