On the Usage of Structural Information in Constrained Semi-Supervised Clustering of XML Documents

Eduardo Augusto Duque Bezerra, Geraldo Bonorino Xexeo, Marta Mattoso · IGI Global eBooks · 2008

In this chapter, we consider the problem of constrained clustering of documents. We focus on documents that present some form of structural information, in which prior knowledge is provided. Such structured data can guide the algorithm to a better clustering model. We consider the existence of a particular form of information to be clustered: textual documents that present a logical structure represented in XML format. Based on this consideration, we present algorithms that take advantage of XML metadata (structural information), thus improving the quality of the generated clustering models. This chapter also addresses the problem of inconsistent constraints and defines algorithms that eliminate inconsistencies, also based on the existence of structural information associated to the XML document collection.

Read the paper · More papers on PaperTik