A Chinese Web page clustering algorithm based on the suffix tree

Jianwu Yang · Wuhan University Journal of Natural Sciences · 2004

In this paper, an improved algorithm, named STC-I, is proposed for Chinese Web page clustering based on Chinese language characteristics, which adopts a new unit choice principle and a novel suffix tree construction policy. The experimental results show that the new algorithm keeps advantages of STC, and is better than STC in precision and speed when they are used to cluster Chinese Web page.

Read the paper · More papers on PaperTik