Structure-based classification of web documents using Support Vector Machine
Kejing He, Chenyang Li · 2016
The web is a huge repository of information and there is a need for web document classification to facilitate the indexing, search and retrieval. Web document classification is significantly different from traditional full text classification because of the existence of some additional information provided by the HTML structure. This paper analyzes the structure information of web documents, and utilizes a structure-based Support Vector Machine (SVM) classifier for classification. The method confirms, quantifies, and extends previous research by introducing a new structure-based method for description and classification of web documents. Compared to traditional web document classification methods, combining the full text with structure information gains nearly 6% accuracy improvement in the case of similar categories and 3.7% accuracy improvement in the case of distinct categories. The structure-based representation of web documents makes use of merely local information, therefore it can be used even in real-time classification.