Optimization of Web Content Mining with an Improved Clustering Algorithm
Sovers Singh Bisht, Sanjeev Bansal · 2013
Abstract—Web Content Mining generally consist of mining the content held within the web pages. The aim of this paper is to evaluate, propose and improve the use of advance web data clustering techniques which is highly used in the advent of mining large content based data sets which allows data analysts to conduct more efficient execution of large scale web data searches. In the search space data is available in a random fashion which may cause trafficking when searched multiple times. Thus in this paper we provide an improved algorithm which may reduce the search space in search engines using clustering techniques. Keywords—Web mining, database, clustering algorithm, web document, data sets, K-mean clustering and Web usage mining. I.