Understanding seed selection in bootstrapping
Yo Ehara, Issei Sato, Hidekazu Oiwa, Hiroshi Nakagawa · 2013
Bootstrapping has recently become the focus of much attention in natural language processing to reduce labeling cost.In bootstrapping, unlabeled instances can be harvested from the initial labeled "seed" set.The selected seed set affects accuracy, but how to select a good seed set is not yet clear.Thus, an "iterative seeding" framework is proposed for bootstrapping to reduce its labeling cost.Our framework iteratively selects the unlabeled instance that has the best "goodness of seed" and labels the unlabeled instance in the seed set.Our framework deepens understanding of this seeding process in bootstrapping by deriving the dual problem.We propose a method called expected model rotation (EMR) that works well on not well-separated data which frequently occur as realistic data.Experimental results show that EMR can select seed sets that provide significantly higher mean reciprocal rank on realistic data than existing naive selection methods or random seed sets.