Identifying Similar Web Pages Based on Automated and User Preference Value Using Scoring Methods

K. Gandhimathi, Vijaya MS · International Journal of Data Mining & Knowledge Management Process · 2013

A World Wide Web (WWW) is a system of interlinked hypertext documents accessed via internet.The web structure mining is based on the graph structure of hyperlinks and it extracts the useful information from structure of web data.Web structure mining aims to generate structural summary about web sites and web pages.Identifying the web community is one of the goal in web structure mining.This paper presents an alternate web community identification approach to detect the similar web pages in web community.The web community identification model is generated by extracting the attributes from the HTML source of input page.In the proposed work the preference value assigned by the user and automatically computed preference value for the given input page are used for analysis.The candidates are identified from the input page.Candidates are scored by three scoring methods such as normalized method, backlink analysis method and hyperlink analysis method and compared with preference value to find the similar pages in web community.The candidate pages with scores equal to preference value of the input pages are recognized as similar pages and hence it is found that the method produces scores equal to preference value which is an enhanced approach for identifying the similar pages.It is verified by using similarity analyzer tool by comparing the candidate pages with input page.

Read the paper · More papers on PaperTik