A Web Site Representation and Mining Algorithm Using the Multiscale Tree Model

Yonghong Tian, Tiejun Huang, Gao Wen · 2015

Abstract: With the exponential growth of both the amount and the diversity of the web information, web site mining is highly desirable for automatically discovering and classifying topic-specific web resources from the World Wide Web. Different with single pages, web sites are essentially heterogeneous, multi-structured and often accompanied with much noise. Nevertheless, existing web site mining methods have not yet handled adequately how to make use of various contextual semantic clues and how to denoise the content of sites effectively so as to obtain better classification accuracy. This paper circumstantiates three issues to be solved for designing an effective and efficient web site mining algorithm, i.e., the sampling size, the analysis granularity and the representation structure of web sites. On the basis, this paper proposes a novel multiscale tree representation model of web sites and presents a multiscale web site mining approach, which contains an HMT-based two-phase classification algorithm, a context-based interscale fusion algorithm, a two-stage text-based denoising procedure and an entropy-base pruning strategy. The proposed model and algorithms may also be used for some related mining tasks

Read the paper · More papers on PaperTik