A Web Document Clustering Algorithm Based on Association Rule

Song Qin-bao · 2002

By grouping similar Web documents into clusters, the search space can be reduced, the search accelerated, and its precision improved. In this paper, a new clustering algorithm is introduced. In the clustering technique, topics are represented according to VSM (vector space model), documents are represented according to topics, and the relation between documents and topics is viewed in a transactional form, each document corresponds to a transaction and each topic corresponds to an item. A frequent item sets can be found by using the association rules discovery algorithm, corresponding documents can be seen as initial clusters. These clusters are merged according to the distance between clusters, or divided according to the strength of connection among documents of a cluster. By real Web documents, experimental results show the algorithm抯 effectiveness and suitability for tackling the overlapping clusters inhered by documents.

Read the paper · More papers on PaperTik