An adaptive neural network approach to hypertext clustering

Natalija Vlajic, HOWARD C. CARD · 2003

The WWW is an online hypertextual collection, and a more sophisticated algorithm for Web page clustering may have to be based on combined term-similarity and hyperlink-similarity measures. It has been observed that nearly all currently employed techniques for document classification on the Web make use of textual information only. In addition, most of these techniques are incapable of discovering the real nature of the collection to which they are applied due to rather inefficient clustering algorithms employed. This paper describes a novel technique for hypertext clustering, called an adaptive hypertext clustering (AHC) algorithm. This algorithm has been derived from a modified neural network algorithm, and adjusted to the problem of combined term-similarity and hyperlink-similarity measures. The results presented in the paper show that AHC can be easily adapted to enable the most appropriate Web page classification within collections of various thematic and functional profiles, suggesting its main benefits over the traditional techniques.

Read the paper · More papers on PaperTik