Web Page Classification based on Document Structure
Arul Prakash Asirvatham, K. N. Anjan Kumar · 2001
The web is a huge repository of information and there is a need for categorizing web documents to facilitate the search and retrieval of pages. Existing algorithms rely solely on the text content of the web pages for classification. However, the web has a lot of information contained in structure, images, video etc present in the document. In this paper, we propose a method for automatic classification of web pages into a few broad categories based on the structure of the web document and the characteristics of images present in it.