A Density-based Approach for Positive and Unlabeled Learning
Cholwich Nattee, Thanaruk Theeramunkong, Masayuki Numao · 2012
or impossible to prepare a set of negative examples. We propose a technique to automatically extract a set of reliable negative examples from the collected unlabeled examples. The proposed technique is based on the idea that an unlabeled example has a high chance to be a negative example when it is located far from any positive examples and in an area with various unlabeled examples. From the extracted negative examples, we can then use an ordinary supervised machine learning technique to build a classication model. To evaluate the proposed technique, we conduct experiments based on sets of generated examples. The experimental results conrm the eectiveness of the technique. models from only positive and unlabeled examples without negative examples. This learning technique plays an important role in the application domains where labeling negative examples requires a lot of eort or costs. It is very dicult or impossible to label negative examples in some applications, e.g. text categorization where one example may belong to multiple classes, and it is usually trivial for users to specify the classes that the example belongs to, but it is dicult to specify the class that the example does not belong to. A basic approach to handle this problem of labeling negative examples by using an ordinary supervised machine learning technique is to construct a classier by using the examples in the target class as positive examples, and the examples in the other classes as negative examples.