An algorithm for automatic web-page clustering using link structures

Debajyoti Mukhopadhyay, Sanasam Ranbir Sing · 2005

Web contains a large collection of heterogeneous documents. As a result, finding set of related pages from Web is currently facing one of the most crucial problems. The low precision Web search engines like Excite, Alta Vista etc. coupled with the ranked list presentation make it harder for users to find the information they are looking for. In this paper, we have proposed a methodology to cluster related pages using co-citations without manual study and/or predefined categories. These clusters are used to classify random pages in the Universe.

Read the paper · More papers on PaperTik