Pinpoint Clustering of Web Pages and Mining Implicit Crossover Concepts
Makoto Haraguchi, Yoshiaki Okubo · InTech eBooks · 2010
In this chapter, we presented our Top-N methods for extracting clusters of Web pages, especially, a method for pinpoint clustering of Web pages by pseudo-clique search and a method for finding implicit page groups represented as formal concepts. In our pinpoint clustering, we first extract semantic correlations among terms by applying SVD to the term-document matrix generated from a corpus w.r.t. a specific topic. Based on the correlations, we can evaluate potential similarities among Web pages from which we try to obtain clusters. The set of Web pages is represented as a weighted graph G based on the similarities and their ranks. Then our clusters are extracted as pseudo-cliques in G. Our experimental results showed that a valuable cluster can be actually extracted according to our method. Turning our attention from clique-based clusters to formal concept-based clusters in order to make our clusters more meaningful, we discussed an effective depth-first mining algorithm for finding relatively smaller therefore more implicit groups of Web pages as formal concepts. The algorithm is based on a dynamic ordering method depending on each search node and some search tree expansion rules. Moreover it was designed so as to find Top-N implicit concepts subject to the size restriction and some space constraints reflecting user's interests. Our experimental results showed that our Top-N algorithm succeeds in finding less frequent (crossover) concepts under some space constraints.