Clustering GML documents using maximal frequent induced subtrees

Ying-wen Zhu, Genlin Ji, Qin-hong Sun · 2010 Seventh International Conference on Fuzzy Systems and Knowledge Discovery · 2010

An algorithm, TBCClustering, is presented in the paper for clustering GML documents using maximal frequent induced subtree patterns. TBCClustering mines the maximal frequent induced subtrees by using the structural information of GML documents, it can get the best minimum support automatically, and then chooses a set of subtree patterns to form the optimistic clustering features. Finally it uses CLOPE algorithm to cluster the GML documents by clustering features without giving the number of clusters. Experiment results have shown that TBCClustering is more effective and efficient than PBClustering.

Read the paper · More papers on PaperTik