Web Text Clustering Based on Concept Lattice

Yimin Shi, Jun Zhang, Xianzhong Zhang, Yanxia Li · 2010

Most web text clustering is based on the space vector text representation model. This results in a high dimension in the terms; and it leads to an increase in time complexity and a loss of text semantics due to the fact that the semantic relationship of the terms is not considered. In this paper, a new approach is taken where a concept lattice is generated with text treated as object and terms of text as attribute to construct a concept lattice. Based on this, formal concepts in the concept lattice are extracted to represent the texts. In addition, similarity function between concepts is defined. To address the drawbacks of the existing K-Means algorithm, such as random selection of initial center, a method is proposed which takes into account the density and distance factors comprehensively. This new algorithm has been applied to the clustering module of our existing maritime vertical searching engine "Haisou". The results demonstrate improved clustering efficiency and accuracy.

Read the paper · More papers on PaperTik