Enhancing Contents-Link Coupled Web Page Clustering and Its Evaluation
Yitong Wang, Masaru Kitsuregawa · 2004
Web page clustering is a fundamental technique to offer a solution for data management, information locating and its interpretation of Web data and to facilitate users for navigation, discrimination and understanding. Most existing clustering algorithms cannot adapt well to Web clustering directly in terms of efficiency and effectiveness. Combining contents analysis and hyperlink structure analysis has been proven a better approach. However, how to effectively combine the two features with different nature in clustering to get satisfactory results remains an open problem and there is still little work on it. In this paper, we present an experimental study on enhancing coupling of links and contents analysis of Web pages for robust clustering. In particular, we introduce two techniques: in-link reinforcement and anchor window analysis to improve the adaptability of contents-link coupled clustering. Our detailed evaluation indicates those techniques can effectively improve the quality of Web pages clustering for a wide range of topics. 1. Introduction there are more than 2 billion pages on the web without counting those so-called hidden Web pages that can be generated from the underneath databases. At the same time more than 100 million pages become obsolete every month. Locating truly needed Web pages and interpreting them appropriately is a big challenge faced by researchers in the fields of database, Information Retrieval (IR) and data mining. So, correctly clustering both the source Web pages and results of search engines is very important to help end users in navigation, discrimination, summarization and interpretation of the Web. Most existing and well-cited topic directories such as Yahoo! (www.yahoo.com) and open directory (www.dmoz.com) are mainly created and maintained manually by domain experts. Therefore those topic directories cover only a very small portion of the whole Web due to extremely low scalability of manual creating and maintenance. They are also more often outdated as the Web changes all the time. Some topics also have no corresponding sub-categories in Yahoo or open directory. Such unsatisfactory performance calls for the needs of semi-automatic or automatic clustering of Web pages that is expected to scale well and be able to follow the evolution of the Web well. Document clustering has been well studied in the field of tradition IR. The most commonly used techniques are developed under the vector-space model. Under this