A Similarity-based Soft Clustering Algorithm for Web Documents

Zequn Guan · Jisuanji gongcheng · 2006

This paper proposes similarity-based soft clustering (SISC), an efficient soft clustering algorithm based on a given similarity measure used in document clustering. Comparison with existing hard clustering algorithms like K-means, the experiment indicates SISC is both efficient and effective, and this algorithm is available for document clustering. In the end, it highlights the upcoming challenges of document mining and the opportunities it offers.

Read the paper · More papers on PaperTik