On Supervised Clustering for Creating Categorization Segmentations
Stephen C. Gates, Charų C. Aggarwal, Arindam Banerjee · Chapman & Hall/CRC data mining and knowledge discovery series · 2008
In this paper, we discuss the merits of using supervised cluster- ing for coherent categorization modeling. Traditional approaches for docu- ment classification on a predefined set of classes are often unable to provide sufficient accuracy because of the difficulty of fitting a manually categorized collection of records in a given classification model. This is especially the case for domains such as text in which heterogeneous collections of Web documents have varying styles, vocabulary, and authorship. Hence, this paper investi- gates the use of clustering in order to create the set of categories and its use for classification. We will examine this problem from the perspective of text data. Completely unsupervised clustering has the disadvantage that it has difficulty in isolating sufficiently fine-grained classes of documents relating to a coherent subject matter. In this chapter, we use the information from a pre-existing taxonomy in order to supervise the creation of a set of related clusters, though with some freedom in defining and creating the classes. We show that the advantage of using supervised clustering is that it is possible to have some control over the range of subjects that one would like the cat- egorization system to address, but with a precise mathematical definition of how each category is defined. An extremely effective way then to categorize documents is to use this a priori knowledge of the definition of each category. We also discuss a new technique to help the classifier distinguish better among