Selective Domain Information Acquisition to Improve Segmentation Quality

Yinghui Yang, Zijie Qi, Hongyan Liu · 2015

It has been well established that adding domain information about whether certain data objects (for example customers in customer segmentation application) should belong to the same segment can improve the quality of the segments. However, it can be expensive to acquire such domain knowledge. Consequently, we need to limit the number of constrains we acquire and more importantly maximize the effectiveness of these limited number of constraints we can acquire from the experts. Many of the constrained clustering methods randomly select constraints, which have been shown in our experiments to be ineffective. In this paper, we define the problem of identifying the most informative constraints. We propose an algorithm which generates the most informative constraints by maximizing the information gain from the constraints. We conducted a set of experiments on various data sets to compare our method with two other methods according to two measurements: the accuracy rate and the Vector Quantization Error (VQE), which is the objective function k-means clustering method minimizes. We illustrated that our approach not only achieves better accuracy rates, but also maintains low VQE values. Our results suggest that businesses can enhance their segmentation quality greatly by actively acquire the right type of domain information.

Read the paper · More papers on PaperTik