Relation Extraction and Discovery from Free Texts via Bootstrapping

Yunlong Yang, Jie Luo · 2017

Considering the problem of extracting binary relations from free texts for different domains, we use a universal method with no human supervision. Previous methods based on machine learning tended to extracting relations for a specific domain. When applying to different domains, these methods needed to be retrained based on new features and large amounts of manually labeled data. In this paper, we propose a bootstrapping method of relation extraction which can be adapted to different domains without training and utilizing of domain specific features. Firstly, a set of seeds is generated and the contexts of sentences which contain seeds are extracted as 5-tuples. Secondly, 5-tuples are clustered based on a small set of domain independent features. Thirdly, the 5-tuple whose feature vector has the highest similarity with the central vector of each cluster is chosen as a candidate pattern. Finally, all candidate patterns are evaluated and only the reliable patterns are selected for extracting new relations which shall be added to the set of seeds in next iteration. The experimental results show that our approach outperforms several other methods.

Read the paper · More papers on PaperTik